On this page9 sections
How accurately can AI read a UK receipt or invoice? Every receipt-scanning app says "very". Few say how they know, what they tested on, or where it goes wrong.
So we published ours. On 28 September 2026 we ran 116 UK documents through the exact code TapTax runs in production to read receipts and bills, scored every field against what is printed on the page, and wrote down every failure. This post is the result: what we tested, how we scored it, where it worked and where it did not.
What did we test, and why synthetic documents?
116
documents in the corpus
99.8%
core fields read correctly
95.6%
line items read correctly
The corpus is 116 documents, every one generated by a seeded program rather than collected from customers. That was a deliberate choice: a benchmark we can publish in full cannot contain anyone's real financial records. The generator produces documents the way they arrive in a sole trader's life.
| Format | What it is | Documents |
|---|---|---|
| Text PDF | A digital PDF with a text layer, as accounting software sends it | 45 |
| Scanned PDF | An image-only PDF: greyscale, skewed and speckled | 15 |
| Multi-page text PDF | Two or three pages with a text layer | 4 |
| Multi-page scanned PDF | Two image-only pages | 2 |
| Photo | A paper invoice photographed at an angle, some in poor light | 24 |
| Till receipt photo | A thermal receipt photographed on a surface | 20 |
| Handwritten | A receipt-book page written by hand | 6 |
By type, that is 76 invoices (including Construction Industry Scheme and domestic reverse charge invoices), 26 receipts, 8 credit notes, and 6 documents that are not financial at all: a letter, a menu, a flyer, a delivery note with no prices, a parking notice and some meeting notes. Those six are there to check the pipeline knows when to stop.
How did we decide what counted as correct?
- Invented value
- A value the reader returned for a field the document does not print. We count these separately, as errors, rather than ignore them.
The rule that matters most: labels record what is printed on the page, never what "should" be there. A field is only scored where the page prints it, and a value returned for a field the page does not print is counted against us as invented.
Beyond that, the matching rules were strict:
- Money and VAT rates had to be equal to the penny.
- Dates had to be exactly the right day.
- Supplier names had to match after lower-casing, dropping punctuation and a trailing Ltd, Limited, LLP, PLC or Co.
- VAT numbers had to have the same digits, with or without spaces and the GB prefix.
- Invoice numbers and references had to have the same letters and digits, ignoring spaces and punctuation.
- A line item counted only when its description matched closely and every column the document prints, quantity, unit price and amount, matched to the penny.
The pass marks were set on 23 August 2026, a month before the run: 95% for core fields and 85% for line items. Setting them first meant they could not be quietly fitted to whatever the result turned out to be.
What were the results?
| Measure | Result | Bar set in advance | Verdict |
|---|---|---|---|
| Core fields (supplier, invoice number, date, net, VAT, total) | 625 of 626, 99.8% | 95% | Met |
| Line items | 481 of 503, 95.6% | 85% | Met |
| Document type | 116 of 116, 100% | none | n/a |
| Cost to read | 2.61p a document | none | n/a |
Every one of the 16 fields we score was read correctly every time it was printed, except one: the invoice number, which was right 103 times out of 104. The supplier name, address, VAT number, invoice date, net, VAT and total were right on every document that printed them.
All six non-financial documents were recognised as such and turned away, at a fifth of a penny each, before any detailed reading was spent on them.
Where did it go wrong?
This is the part most benchmarks leave out.
Till receipt line items: 78.9%. On photographed thermal receipts, about four lines in five were read with every column right. The header and the total on the same receipts were read perfectly, so the amount that reaches your records was correct; it is the line-by-line breakdown that suffers. Thermal print is small, low in contrast and often crumpled, and that shows.
Scanned PDF line items: 84.8%. Just under the 85% bar for that format on its own, though the corpus as a whole clears it. Core fields on the same scans were 100%.
Ten invented values across 116 documents. A VAT rate was returned seven times on documents that print none, and an invoice number three times on documents that print none. We count those against the pipeline. In the app, each field carries a confidence and anything doubtful is put in front of the user to check, which is what catches this kind of error before it reaches a tax figure.
Two extra line items were returned that the pages do not print.
“The totals were right on every till receipt. The line-by-line breakdown was right four times in five. Both of those are true, and you need both to judge a receipt scanner.”
How did each kind of document do?
| Format | Core fields | Line items | Mean cost |
|---|---|---|---|
| Text PDF | 100.0% | 100.0% | 1.39p |
| Multi-page text PDF | 100.0% | 100.0% | 3.10p |
| Handwritten | 100.0% | 100.0% | 3.53p |
| Multi-page scanned PDF | 100.0% | 100.0% | 11.38p |
| Scanned PDF | 100.0% | 84.8% | 4.27p |
| Photo | 99.2% | 94.4% | 2.68p |
| Till receipt photo | 100.0% | 78.9% | 2.78p |
The pattern is the one you would expect. A PDF straight from accounting software reads perfectly and cheaply, because the text is already there. Images cost more to read and lose accuracy on fine detail. The handwritten receipts did better than we expected, though six documents is too few to draw a firm conclusion.
By document type, invoices scored 99.8% on core fields and 97.6% on line items, receipts 100% and 81.0%, and credit notes 100% on both.
How does the reading actually work?
Each document is first classified from its first page or a small preview: invoice, receipt, credit note, or not financial. A PDF with a text layer is then read by a fast, low-cost model; an image or scan goes to a stronger vision model; and any field that comes back doubtful can be escalated to the most capable model. The models are Anthropic's Claude family, which our privacy policy names as the processor.
After reading, every figure is checked by ordinary code, not by the model: net plus VAT must equal the total, the VAT rate must be a UK rate, and the dates must make sense. A document that fails a check is flagged rather than corrected by guesswork. A transient failure is retried up to three times, as the production job system does, and every retry is included in the cost.
What are the limits of this benchmark?
- The documents are synthetic. That protects real people's data, but real receipts are messier than any generator: torn, folded, faded, photographed in a van at dusk. Expect real-world accuracy to be lower, most of all for photographed receipts.
- 116 documents shows the shape, not a margin of error. Groups as small as six handwritten receipts or two multi-page scans are indicative only.
- It measures reading, not judgement. Whether a cost is allowable, or which HMRC expense category it belongs in, is not scored here. That is a separate question, covered in our guide to what expenses you can claim.
- It measures our pipeline. Other products may do better or worse; we have not tested them, and vendor accuracy claims we quote elsewhere are theirs.
Why publish this at all?
Under Making Tax Digital, sole traders and landlords with qualifying income over £50,000 keep digital records from 6 April 2026, with the threshold falling to £30,000 in April 2027 and £20,000 in April 2028. For many of them, the receipt scanner is now part of how their tax return is built. A number on a marketing page is not enough to trust that with; a method, a corpus and the failures are.
We intend to re-run the benchmark as the pipeline changes and publish each run, including any that go backwards.
Try it on your own documents
You can run the same reader on your own receipts and invoices, without an account, and see each field and its check:
- The free invoice parser reads one invoice into supplier, dates, VAT, totals and lines.
- The receipt to CSV converter turns a batch of receipts into one spreadsheet.
Neither keeps your file. In TapTax itself, receipts and bills are read, checked, matched to your bank statement and filed with your quarterly update; the receipts and bills page shows how.
People also ask
How accurate is AI at reading receipts?
In our September 2026 benchmark of 116 synthetic UK documents, the reader got 99.8% of core fields (supplier, date, net, VAT, total) and 95.6% of line items right. Photographed till receipts were the weak spot, at 78.9% for line items, though their totals were all correct.
Is a scanned receipt good enough for HMRC?
HMRC accepts clear digital copies of records. A scanned receipt that shows the supplier, date, amount and any VAT is a valid record, and you must keep business records for at least 5 years after the 31 January deadline for the tax year.
Why test on synthetic documents rather than real ones?
So the whole benchmark can be published without exposing anyone’s financial records. The trade-off is that real documents are messier, so real-world accuracy will be lower, especially for photographed receipts.
What does it cost to read a receipt with AI?
In this run, 2.61p a document on average at list API prices, from 1.39p for a digital PDF to 11.38p for a two-page scan. Documents that turned out not to be financial cost about 0.2p to recognise and set aside.