Classify the errors before you trust the score
A forced model swap in Mousing, nine benchmark runs, and why the expensive model's lead mostly belonged to the answer key

OpenAI is shutting off gpt-5.4-nano on 1 April 2027. In Mousing, my document extraction platform, nano did two jobs: it classified incoming documents, and it was the verifier, the second model that re-reads each page and checks every extracted value.
Swapping the model is one environment variable. Trusting the swap is not. Every number I had for the verifier (how well its confidence separates right from wrong, the review threshold, the token budget) was measured on nano. None of it transfers to a different model.
So I re-benchmarked before choosing, and the first answer was wrong.
The setup
Nine runs through the real pipeline: the production extraction template, the verifier and the plain-code magnitude check, all over the same 100 photographed receipts from the public CORD dataset, on the same day. Seven model pairings, plus two of them again with reasoning effort turned down.
Before reading anything into it, I measured the noise. Two runs with the same extractor scored 0.88 and 0.87 F1 and got 35 and 34 receipts fully correct. So a hundredth of F1 and a few receipts either way is just run-to-run variance.
The first answer
One pairing stood out. With gpt-6.1-sol as the extractor:
| Extractor / verifier | F1 | Receipts fully correct | Cost per page (est.) |
|---|---|---|---|
5.4-mini / 5.4-nano (production) |
0.88 | 35 | ~$0.004 |
6-luna / 6-luna, low effort |
0.89 | 38 | ~$0.0009 |
6.1-sol / 6-luna |
0.91 | 52 | ~$0.010 |
Everything else landed between 33 and 38 fully correct receipts. Sol produced 35% fewer wrong fields than any other arm. Clearly outside the noise, at roughly ten times the cost of the cheapest pairing. I was ready to pick it.
Then I read the wrong answers
Every wrong field from all nine runs, classified:
- 60% were item names, and many were the answer key's fault. CORD's ground truth is built from OCR tokens. The "truth" for one item was
CHOCO CUS ARD PASTRY. A model that wroteCHOCO CUSTARD PASTRYwas marked wrong. - 26% were rows out of line. Items are scored by position. One receipt's labels listed a size modifier as its own item, and that single receipt caused about a fifth of all errors in every run.
- One field was ambiguous. Mousing's receipt template had a single
pricefield. Mini and luna filled it with the unit price. Sol filled it with the line total. CORD's labels use line totals. Part of sol's lead was guessing the same way the answer key did.
Re-scored with fuzzy matching on names, the gap closed: mini plus nano 0.90, luna 0.92, sol 0.92.
The fix for the ambiguity wasn't a model. I split price into unit_price and amount, both "exactly as printed, never calculated", and re-ran three arms:
| mini + nano | luna, both roles, low effort | sol + luna | |
|---|---|---|---|
Line total (amount) F1 |
0.76 | 0.90 | 0.89 |
| Unit price F1 | 0.48 | 0.89 | 0.83 |
| Item name F1 | 0.81 | 0.80 | 0.85 |
| Cost per page (est.) | ~$0.004 | ~$0.0009 | ~$0.010 |
On the money fields, the ones a customer actually acts on, luna matched sol on line totals and beat it on unit prices, at about a tenth of the cost. Sol's remaining lead is spelling item names, which is exactly where the label noise lives. Mini, meanwhile, filed the line total under unit_price on 37 lines, which means the new schema must never ship without the model change.
Two things I'd have missed
Reasoning effort was free. Turning luna's reasoning effort down kept the same accuracy, cut cost by about a third, and took the slow tail of extraction (the 95th percentile) from 64 seconds to 7. Nobody notices an average. Everybody notices a 64-second wait.
Fewer false flags is a trade, not a win. Nano flagged 39% of correct values for human review, roughly five pointless reviews for every real catch. In practice that's "a person reviews everything." The newer verifiers cut false flags by about 95%, and let more real errors through without a flag. That's a product decision about review load against silent errors, and I'd rather make it on purpose than discover it.
What ships
gpt-6-luna for extraction, verification and classification, at low effort. It's in production now.
The one number I'd keep from all of this isn't an F1 score. It's this: the first table pointed at a model that costs ten times more, and the reason was mostly in the benchmark. A score is a summary. The wrong answers are the evidence.
Numbers are one run per arm, on Indonesian café receipts, not the business documents Mousing is built for. Costs are estimates from token counts at OpenAI's standard prices on 2026-10-02.




