1 Sept 2026 · 8 min read · project: docextract
What 35 documents got wrong about arithmetic self-correction
I gave DocExtract's Anthropic backend a tool to check its own arithmetic during extraction instead of catching the error after, like the existing hard rule does. On a 35-document slice it looked like a clean win. Reran on the full 361-document split, the same N this project already treats as its honest sample size, and the single-call ranking flipped entirely, the tool's real edge shrank to about a third of what the small sample suggested, and it introduced a new failure mode of its own.
The question, and where the tool sits
DocExtract's existing pipeline runs a hard rule after extraction: if the subtotal and tax don't add up to the stated total, or the line items don't add up to the subtotal, the document gets routed to manual review no matter how confident the model was. The question I wanted an answer to: would giving the model a tool to check its own arithmetic during extraction let it self-correct and auto-accept documents that currently get punted?
Key decision
validate_arithmetic isn't an LLM grading its own math. It's a client-side recomputation: sum the line items in Python, compare against the stated total, hand the boolean result back as a tool response. The model gets up to 3 rounds to call it, see a mismatch, and revise, the same bounded-retry pattern the existing backends already use for API calls, so a model that can't converge doesn't loop forever.The numbers, at the scale that actually counts
Three backends, same 361-document split, gemini-2.5-flash and claude-haiku-4-5:
| Backend | Auto-accept | Critical P (total) | $/doc | p50 latency |
|---|---|---|---|---|
| gemini | 117/361 (32.4%) | 116/117 (99.1%) | $0.00090 | 8.94s |
| anthropic | 79/361 (21.9%) | 76/79 (96.2%) | $0.00482 | 4.09s |
| anthropic-agentic | 126/361 (34.9%) | 121/126 (96.0%) | $0.01538 | 8.48s |
What actually happened
anthropic (25.7%) auto-accepting more than gemini (22.9%), and the agentic backend recovering 4 documents that anthropic's own extraction had hard-failed on, with no precision cost anywhere on that slice. That looked like a clean answer: yes, self-correction wins, by 8.6 to 11 points depending on which single-call backend you compared it to. At 361 documents the single-call ranking flips outright: gemini auto-accepts 32.4% against anthropic's 21.9%, a 10.5-point gap in the other direction from what the small sample showed. The agentic backend is still the highest auto-accept rate of the three, but the baseline it should be measured against matters: its edge is +13.0 points over anthropic and only +2.5 points over gemini, the backend that turned out to be the actually stronger single-call option. Roughly a third of the apparent gain survived the rerun. The recovery mechanism itself did generalize, 50 documents auto-accepted at 361 where anthropic's own extraction had failed the reconciliation check, up from 4 at small scale, so the tool is doing real work. It just isn't free: 4 of the agentic backend's 5 critical misses share the same new shape, the tool reconciling a total that was already correct into a self-consistent but wrong number. On X51006328967 (gold total 62.00), the plain backend read total correctly as 62.00 but flagged it for review because its own subtotal and tax readings didn't sum to it. Given the tool, the agentic backend didn't re-read the source for the actual misread field, it adjusted total to 65.51 until the arithmetic closed, and auto-accepted a document that was right before it touched it. That's the exact kind of confidently-wrong auto-accept this whole project's precision posture is built to catch, produced here by the self-correction mechanism itself.Cost and latency were the one thing that didn't move between the two runs. The agentic backend's extra self-correction rounds cost ~3.2x anthropic's per-document price ($0.01538 vs $0.00482) and ~2.1x its p50 latency (8.48s vs 4.09s), both from the same model so price and speed aren't confounded by a model swap. That's close to the 35-document slice's ~3.07x/~2x figures. Running the full split also surfaced a real bug the small slice never hit: the arithmetic tool crashed on a document with a null line-item amount, something that only ever showed up once there were enough documents to hit it. Fixed and re-predicted before any of the numbers above were run.
Known limitations
Stated plainly rather than left for someone else to find:
- It isn't strictly monotonic. 9 documents anthropic had auto-accepted, the agentic backend sent to review instead. 8 of those are true regressions: the plain backend's total was already correct, and self-correction volunteered extra detail that broke a consistency check the shorter answer had passed. The 9th wasn't a regression, anthropic's own total was already wrong against gold and the agentic backend correctly declined to accept it.
- SROIE only labels total as a critical field. This comparison can't speak to precision on tax or invoice_number at all, gold labels for those don't exist in the dataset, so the precision numbers above are total-only by construction, not by choice.
- n=361 is the full SROIE test split, not a larger benchmark. The single-digit miss counts per backend (1, 3, 5) mean a percentage point of critical precision here is worth roughly 3 to 4 documents. Enough to trust the direction of the numbers, not enough to treat the gap between backends as settled.
- The cost/latency tradeoff is the one number I'd act on today. It held steady across a 10x change in sample size while the accept-rate and precision numbers it's traded against didn't. Whether a few points of auto-accept rate is worth 3x the per-document cost, plus a new failure mode, is a product decision this finding doesn't make for you.
The reason this is worth writing up isn't the tool, it's that a 35-document eval looked completely conclusive and reported the wrong winner. A sample size that's fine for a gut check isn't the same thing as a result, and the only way I actually found that out was rerunning at the N I already trust everywhere else in this project.
Python · Anthropic API · Gemini API · Forced tool-use · ICDAR SROIE · pytest