1 Sept 2026 · 8 min read · project: docextract

What 35 documents got wrong about arithmetic self-correction

I gave DocExtract's Anthropic backend a tool to check its own arithmetic during extraction instead of catching the error after, like the existing hard rule does. On a 35-document slice it looked like a clean win. Reran on the full 361-document split, the same N this project already treats as its honest sample size, and the single-call ranking flipped entirely, the tool's real edge shrank to about a third of what the small sample suggested, and it introduced a new failure mode of its own.

Boundary, stated up front: the first version of this finding ran on 35 documents and reported the wrong winner. Everything below is the corrected run on the full 361-document SROIE test split, the same N this project already treats as its honest sample size everywhere else it reports numbers. Even that isn't huge: the miss counts behind the precision numbers here are 1, 3, and 5 documents per backend, so read the precision deltas between backends as directional, not as statistically separated from each other.
361
Documents in the real run, not 35
+2.5
Agentic's real edge, in points, over the actual best single-call backend
3.2x
Cost per document vs. plain Anthropic
4/5
Agentic misses from one new failure mode

The question, and where the tool sits

DocExtract's existing pipeline runs a hard rule after extraction: if the subtotal and tax don't add up to the stated total, or the line items don't add up to the subtotal, the document gets routed to manual review no matter how confident the model was. The question I wanted an answer to: would giving the model a tool to check its own arithmetic during extraction let it self-correct and auto-accept documents that currently get punted?

00
Extract
forced tool-use, schema-constrained
→
01
validate_arithmetic
agentic backend only, up to 3 rounds
→
02
Reconcile
subtotal + tax vs. total, line items vs. subtotal
→
03
Route
auto-accept or review

Key decision

validate_arithmetic isn't an LLM grading its own math. It's a client-side recomputation: sum the line items in Python, compare against the stated total, hand the boolean result back as a tool response. The model gets up to 3 rounds to call it, see a mismatch, and revise, the same bounded-retry pattern the existing backends already use for API calls, so a model that can't converge doesn't loop forever.

The numbers, at the scale that actually counts

Three backends, same 361-document split, gemini-2.5-flash and claude-haiku-4-5:

BackendAuto-acceptCritical P (total)$/docp50 latency
gemini117/361 (32.4%)116/117 (99.1%)$0.000908.94s
anthropic79/361 (21.9%)76/79 (96.2%)$0.004824.09s
anthropic-agentic126/361 (34.9%)121/126 (96.0%)$0.015388.48s

What actually happened

The 35-document slice had anthropic (25.7%) auto-accepting more than gemini (22.9%), and the agentic backend recovering 4 documents that anthropic's own extraction had hard-failed on, with no precision cost anywhere on that slice. That looked like a clean answer: yes, self-correction wins, by 8.6 to 11 points depending on which single-call backend you compared it to. At 361 documents the single-call ranking flips outright: gemini auto-accepts 32.4% against anthropic's 21.9%, a 10.5-point gap in the other direction from what the small sample showed. The agentic backend is still the highest auto-accept rate of the three, but the baseline it should be measured against matters: its edge is +13.0 points over anthropic and only +2.5 points over gemini, the backend that turned out to be the actually stronger single-call option. Roughly a third of the apparent gain survived the rerun. The recovery mechanism itself did generalize, 50 documents auto-accepted at 361 where anthropic's own extraction had failed the reconciliation check, up from 4 at small scale, so the tool is doing real work. It just isn't free: 4 of the agentic backend's 5 critical misses share the same new shape, the tool reconciling a total that was already correct into a self-consistent but wrong number. On X51006328967 (gold total 62.00), the plain backend read total correctly as 62.00 but flagged it for review because its own subtotal and tax readings didn't sum to it. Given the tool, the agentic backend didn't re-read the source for the actual misread field, it adjusted total to 65.51 until the arithmetic closed, and auto-accepted a document that was right before it touched it. That's the exact kind of confidently-wrong auto-accept this whole project's precision posture is built to catch, produced here by the self-correction mechanism itself.

Cost and latency were the one thing that didn't move between the two runs. The agentic backend's extra self-correction rounds cost ~3.2x anthropic's per-document price ($0.01538 vs $0.00482) and ~2.1x its p50 latency (8.48s vs 4.09s), both from the same model so price and speed aren't confounded by a model swap. That's close to the 35-document slice's ~3.07x/~2x figures. Running the full split also surfaced a real bug the small slice never hit: the arithmetic tool crashed on a document with a null line-item amount, something that only ever showed up once there were enough documents to hit it. Fixed and re-predicted before any of the numbers above were run.

Known limitations

Stated plainly rather than left for someone else to find:

The reason this is worth writing up isn't the tool, it's that a 35-document eval looked completely conclusive and reported the wrong winner. A sample size that's fine for a gut check isn't the same thing as a result, and the only way I actually found that out was rerunning at the N I already trust everywhere else in this project.

Python · Anthropic API · Gemini API · Forced tool-use · ICDAR SROIE · pytest