Our Lead Agent demo prepares a discovery brief for a fictional prospect. The prospect reconciles 3,000 subscriptions a week between HubSpot and Stripe, found 240 discrepancies last week and wants a read-only discrepancy report for an operations lead.

A good brief prepares the first conversation. It does not solve the reconciliation. To make that boundary concrete, we built the reconciliation itself as a separate checked example. It makes no model calls.

Everything below uses a constructed dataset with known answers. There is no live HubSpot or Stripe connection.

Write the comparison contract first

Most of the difficulty in reconciliation is deciding what "the same" means. The example states its contract explicitly:

  • rows join on one stable subscription ID, exactly, with no fuzzy matching;
  • both sides must describe the same monthly period;
  • amounts are EUR, as non-negative safe integer cents;
  • statuses come from one agreed vocabulary: active, paused, cancelled, past_due.

These are normalized comparison fields, not raw provider schemas. For a real system, someone has to decide whether an amount includes tax, discounts, proration and refunds, and how each tool's statuses map to the shared vocabulary. That decision belongs to the customer's finance or operations owner, not to a model.

Every pair gets one outcome and its reasons

The engine indexes both sides by subscription ID and gives every ID one of three outcomes:

Outcome When Amount delta
match Both rows valid, same amount and status 0
discrepancy Amount or status differs, or one side is missing Known, or null when a side is missing
review Duplicate ID, period mismatch, unsupported currency, invalid amount, unknown status null

The core of the comparison is short. Simplified from the demo:

if (reasons.length) outcome = "review";
else {
  amountDeltaCents = billing.amountCents - crm.amountCents;
  if (!Number.isSafeInteger(totalDeltaCents + amountDeltaCents))
    throw new Error("Amount total exceeds safe integer range");
  if (amountDeltaCents) reasons.push("amount_mismatch");
  if (crm.status !== billing.status) reasons.push("status_mismatch");
  outcome = reasons.length ? "discrepancy" : "match";
}

Two choices matter here. A missing counterpart is flagged as a discrepancy, but its monetary value stays unknown: we do not invent the amount of a row that is not there. And if the running total would leave the safe integer range, the run stops instead of rounding.

Each outcome keeps its source rows and reason codes, so a reviewer can see why a pair was flagged without re-running anything.

Check against answers the engine never sees

A reconciliation that reports 240 discrepancies looks right when you expected 240. That is not a test. We needed an independent answer for every row.

The generator produces 3,000 CRM rows, 3,000 billing rows and a separate expected-result manifest:

Expected case Count
Matching subscriptions 2,730
Amount differences of +€5.00 180
Status differences 60
Unsupported currency (USD), review 30

The engine does not read the manifest. A verifier compares every pair afterwards: classification, reason codes and per-pair amount delta, plus the supported total.

On this dataset, the run classifies all 3,000 pairs as expected. It finds all 240 planted discrepancies with no false positives and no missed cases, and every reason code matches. The supported amount delta is €900.00, with zero cents of error. The comparison itself runs in milliseconds; the build regenerates and re-verifies it each time.

What the numbers mean

The €900 is an amount difference to investigate. It is not recovered revenue or a customer saving.

It also does not cover everything. The 30 USD pairs remain unresolved in the review queue, excluded from the supported total, because the contract has no currency conversion rule. Converting them at some assumed rate would produce a tidier total and a less honest one.

The public page publishes the summary, four source-linked examples, the 30-item review queue with pending decisions, and an evidence file with every input row, expected case and actual outcome. Anyone can check the classifications without our source code.

Why no model?

Joining on an exact ID and subtracting integers are tasks where a model adds cost, latency and a new failure mode without adding judgment. We use models elsewhere in the same story: a decision model selects discovery context, and a generative model drafts the Lead brief. Neither is asked to certify financial arithmetic.

Where a model might help later is in the work around the comparison: explaining a cluster of review cases, or drafting questions for the person who owns the status mapping. That would be a separate, measured addition, not a replacement for the deterministic check.

What this does and does not establish

Zero errors here establishes that this implementation passes its own controlled check on a constructed dataset with declared rules. It does not establish production accuracy, robust entity matching, correct interpretation of a customer's rules, reviewer acceptance, reduced investigation time or savings.

For a real pilot, the useful promise is a process:

  1. Agree the comparison rules with the people who own them.
  2. Check the implementation against an approved reference set, including representative exceptions.
  3. Keep ambiguous cases visible instead of forcing them into a total.
  4. Measure how much investigation work remains for the team.

Only after those steps can anyone make a correctness or economic claim.

You can inspect the reconciliation example and download its evidence. If your team reconciles records between two systems every week, a sample export and your current matching rules are a good starting point for a pilot conversation.