Finance workflow · Agent demo

Revenue rose.
Where did the margin go?

Follow a margin decline from the overall change to its source records. Explore a large SAP-shaped demo dataset, then a smaller reconciled example with questions for a finance reviewer.

This is a local demonstration using invented data, not a client result. The agent used read-only tools; its interpretation and proposed actions have not been approved by a human reviewer.

Large-data investigation · SAP-shaped demo exports

600,000 rows. Two adverse drivers.

Revenue grew from €95,000,000 to €102,850,000, while gross profit fell by €1,450,000. The investigation follows the change to allocated costs and discount records.

SAP already provides margin analysis and drilldowns. This example shows a bounded agent preparing evidence for a reviewer using saved exports. All data is invented; these are custom demo export projections, not validated SAP API responses.

Billed positions

200,000

Plus one pricing and one cost record per position; 1,000 materials.

Data processing

2.48 s

Parse, join, validate, hash and aggregate 91.3 MB of local exports.

Full live run

15.2 s

Includes generation, both Jev decisions and the LLM draft.

Jev context

1.4–1.5 KB

Per decision request. Two calls cost $0.00005237; full workflow cost is unavailable.

Process / demo

Large exports, bounded model decisions

Observed live Jev and LLM run · SAP-shaped demo data

  1. Input

    Billing, pricing and allocated costs

    Generated JSONL exports represent two periods and one thousand materials. No SAP system is connected.

  2. Deterministic code

    Validate every row and reconcile the change

    Code joins document/item keys, checks currencies and amounts, hashes the full input and calculates the profit bridge in cents.

  3. Model decision

    Jev chooses the next evidence topic

    Compact semantic state lets Jev select cost evidence first, then discount evidence. Code controls allowed topics and review paths.

  4. Read-only tool

    Inspect bounded source examples

    Each lookup returns aggregates covering all affected items and at most eight representative billed positions. The full dataset stays outside model context.

  5. Evidence and validation

    Draft from checked facts

    LLM prepares a provisional explanation and document references. Financial values come from code; citations must match inspected evidence.

  6. Human review

    Review allocations and discount terms

    A finance reviewer checks the accounting assumptions and business causes. No prices, postings or payments are changed.

Boundary: local demo evidence and a provisional draft. No SAP connection, writeback or external financial action occurred; human approval remains pending.

The measured change

Kept-unit volume / mix
€3,500,000
List price
€0
Discounts
-€1,650,000
Allocated costs
-€3,300,000
Total gross-profit change
-€1,450,000

Zero-cent reconciliation residual. Margin moved from 36.84% to 32.62%. The bridge describes contributions; it does not prove business causes.

Provisional explanation

The code-calculated summary shows revenue increased while gross profit and margin rate declined. The bridge identifies higher volume and mix as a positive contributor, offset by higher costs and discounts; list price contributed no change. These are descriptive contributors, not verified underlying causes. The selected evidence provides examples only and is not exhaustive.

The reviewer checks

  • Validate current-period cost changes against supplier invoices, cost records, and approved cost updates.
  • Check current-period discounts against contracts, approvals, and billing rules for the affected materials.
  • Review the full affected item populations and investigate any exceptions before attributing causes.

Jev selected / costs

Allocated cost increase

-€3,300,000 contribution across 100 materials and 20,000 billed positions. Aggregates cover every affected row; the documents below are representative examples.

Inspect MAT0000 source examples

Baseline · 9000000000/000010

Quantity: 8. Net per unit: €86.45. Allocated cost per unit: €54.60.

  • input/billing-items.jsonl#L1
  • input/pricing.jsonl#L1
  • input/costs.jsonl#L1

Current · 9000100000/000010

Quantity: 9. Net per unit: €86.45. Allocated cost per unit: €84.60.

  • input/billing-items.jsonl#L100001
  • input/pricing.jsonl#L100001
  • input/costs.jsonl#L100001

Jev selected / discounts

Discount increase

-€1,650,000 contribution across 100 materials and 20,000 billed positions. Aggregates cover every affected row; the documents below are representative examples.

Inspect MAT0100 source examples

Baseline · 9000000100/000010

Quantity: 8. Net per unit: €86.45. Discount per unit: €4.55.

  • input/billing-items.jsonl#L101
  • input/pricing.jsonl#L101
  • input/costs.jsonl#L101

Current · 9000100100/000010

Quantity: 9. Net per unit: €71.45. Discount per unit: €19.55.

  • input/billing-items.jsonl#L100101
  • input/pricing.jsonl#L100101
  • input/costs.jsonl#L100101
Scaling, routing checks and accounting assumptions

Three isolated Node runs per data size on the same local machine. Processing includes file parsing, joins, validation, hashing and aggregation; it excludes source generation and model calls. Filesystem caching was not controlled.

  • 60,000 export rows (20,000 billed positions): 0.28 seconds mean processing time, three runs.
  • 300,000 export rows (100,000 billed positions): 1.24 seconds mean processing time, three runs.
  • 600,000 export rows (200,000 billed positions): 2.54 seconds mean processing time, three runs.

Each Jev request is under 1.5 KB. LLM receives a 6.2 KB evidence prompt with sixteen representative billed positions, rather than the full export. LLM usage includes runtime context beyond that prompt; API cost for the complete workflow is unavailable.

A separate six-question holdout matched all labels for both Jev and a simple keyword router, with no wrong high-confidence lookups. This is a small test set. Jev adds API latency; these results do not establish a speed or accuracy advantage over rules.

EUR only. Generated billing is fully recognized in its labelled period. No taxes, returns, foreign exchange or material entry/exit are modeled. Pricing and allocated-cost rows are custom projections; a real SAP adapter and recognition mapping need separate validation.

Missing costs, duplicate keys, foreign currency, inconsistent amounts and unsupported document citations block a confident draft. Vague questions and outbound requests require review. No client savings or production throughput are measured.

SAP references: Margin Analysis and Billing Document Item.

A smaller example, step by step.

Process / demo

How the margin investigation works

Recorded demo; interpretation remains a draft

  1. Input

    Business question

    Two fictional monthly periods of Northstar Industries order lines ask why revenue rose while gross profit fell.

  2. Deterministic code

    Calculate the facts

    Code validates the input and reconciles revenue, allocated costs, gross profit and the change bridge.

  3. Model decision

    Choose what to inspect

    The LLM selects a bounded read-only query. The host enforces the tool and turn limits.

  4. Read-only tool

    Read local source lines

    Only query_margin, inspect_orders and inspect_costs can return matching product, order and cost records.

  5. Evidence and validation

    Prepare a cited draft

    The agent writes provisional findings; code rejects unsupported line IDs and non-reconciling drivers.

  6. Human review

    Finance reviewer decides

    A person checks discount terms and cost allocation before accepting any conclusion or action.

Boundary: this run ends with a local draft for human review. It does not change prices, contact suppliers or write to a financial system.

01 / The result

The change, reconciled.

July and August 2026, EUR. Revenue increased, but allocated costs and discounts outweighed the benefit of additional units.

Recognized revenue

+€100

€3,200 → €3,300

Gross profit

−€195

€1,430 → €1,235

Gross-margin rate

−7.26 pp

44.69% → 37.42%

Gross-profit change bridge

A descriptive breakdown of the €195 decline. Returns are included in kept-unit volume and mix.

Kept-unit volume & mix

More FILTER units partly offset the decline.

+€125

List price

No list-price change in this sample.

€0

Discounts

FILTER order line C-FILTER-02.

−€150

Allocated cost of goods

PUMP lines C-PUMP-01 and C-PUMP-02.

−€170

Product entry / exit

The same products appear in both periods.

€0

Total change · €0 rounding residual

−€195

These contributions describe the data; they do not prove why a discount was given or why a cost allocation changed.

02 / The investigation

From question to source lines.

  1. 01 / Verify

    Check the premise

    Code validates the periods and calculates revenue, allocated cost, gross profit and margin. The agent receives those numbers instead of doing arithmetic from memory.

  2. 02 / Follow evidence

    Inspect the lines

    The agent chose bounded read-only queries, then linked the FILTER discount to C-FILTER-02 and the PUMP cost change to C-PUMP-01 and C-PUMP-02.

  3. 03 / Hand off

    Draft review actions

    It proposed checking the FILTER discount against approved terms and reconciling PUMP costs with invoices and allocation records. Those checks remain for a person to perform.

The follow-up question

“Which order lines support the largest cost and discount changes?” The draft answered with source IDs, but did not claim to know whether the original discount or cost allocation was correct.

03 / What we measured

A local run, with its limits.

One inspected live LLM run on 26 September 2026 used gpt-6-luna. It produced a draft in 42.4 seconds with five model turns and three local read-only tool calls. Its arithmetic and source references passed the checks. Three offline replays also passed those checks; replay does not measure current model reliability.

What the demo demonstrates

  • A gross-profit bridge that adds to the measured change.
  • Order-line references for the main discount and cost contributions.
  • A reviewable draft that keeps uncertainty and next checks visible.

What remains unmeasured

  • Human reviewer corrections, time saved and customer outcomes.
  • Production integration reliability and behavior on real financial data.
  • Per-run cost under authenticated LLM; no API price was inferred.

In a separate adversarial run, the agent hit its six-turn limit and returned a blocked result. Its recorded tool calls remained local reads. One sample cannot establish general prompt-injection resistance.

Want to investigate your margin?

We can scope a pilot around your metric definitions, source quality and review process. Real data access and any financial-system action would require a separate design and approval.

Discuss your workflow →