Skip to main content
TapTax

We Tested AI on 116 UK Receipts and Invoices: 99.8% of Core Fields Right, Till Receipts the Weak Spot

An open benchmark of AI document reading on UK receipts, invoices and credit notes: 99.8% of core fields and 95.6% of line items read correctly, bars committed in advance, and every failure mode published.

By Solomon Amos

10 min read

On this page9 sections

How accurately can AI read a UK receipt or invoice? Every receipt-scanning app says "very". Few say how they know, what they tested on, or where it goes wrong.

So we published ours. On 28 September 2026 we ran 116 UK documents through the exact code TapTax runs in production to read receipts and bills, scored every field against what is printed on the page, and wrote down every failure. This post is the result: what we tested, how we scored it, where it worked and where it did not.

What did we test, and why synthetic documents?

116

documents in the corpus

99.8%

core fields read correctly

95.6%

line items read correctly

The corpus is 116 documents, every one generated by a seeded program rather than collected from customers. That was a deliberate choice: a benchmark we can publish in full cannot contain anyone's real financial records. The generator produces documents the way they arrive in a sole trader's life.

FormatWhat it isDocuments
Text PDFA digital PDF with a text layer, as accounting software sends it45
Scanned PDFAn image-only PDF: greyscale, skewed and speckled15
Multi-page text PDFTwo or three pages with a text layer4
Multi-page scanned PDFTwo image-only pages2
PhotoA paper invoice photographed at an angle, some in poor light24
Till receipt photoA thermal receipt photographed on a surface20
HandwrittenA receipt-book page written by hand6

By type, that is 76 invoices (including Construction Industry Scheme and domestic reverse charge invoices), 26 receipts, 8 credit notes, and 6 documents that are not financial at all: a letter, a menu, a flyer, a delivery note with no prices, a parking notice and some meeting notes. Those six are there to check the pipeline knows when to stop.

How did we decide what counted as correct?

Invented value
A value the reader returned for a field the document does not print. We count these separately, as errors, rather than ignore them.

The rule that matters most: labels record what is printed on the page, never what "should" be there. A field is only scored where the page prints it, and a value returned for a field the page does not print is counted against us as invented.

Beyond that, the matching rules were strict:

  • Money and VAT rates had to be equal to the penny.
  • Dates had to be exactly the right day.
  • Supplier names had to match after lower-casing, dropping punctuation and a trailing Ltd, Limited, LLP, PLC or Co.
  • VAT numbers had to have the same digits, with or without spaces and the GB prefix.
  • Invoice numbers and references had to have the same letters and digits, ignoring spaces and punctuation.
  • A line item counted only when its description matched closely and every column the document prints, quantity, unit price and amount, matched to the penny.

The pass marks were set on 23 August 2026, a month before the run: 95% for core fields and 85% for line items. Setting them first meant they could not be quietly fitted to whatever the result turned out to be.

What were the results?

MeasureResultBar set in advanceVerdict
Core fields (supplier, invoice number, date, net, VAT, total)625 of 626, 99.8%95%Met
Line items481 of 503, 95.6%85%Met
Document type116 of 116, 100%nonen/a
Cost to read2.61p a documentnonen/a

Every one of the 16 fields we score was read correctly every time it was printed, except one: the invoice number, which was right 103 times out of 104. The supplier name, address, VAT number, invoice date, net, VAT and total were right on every document that printed them.

All six non-financial documents were recognised as such and turned away, at a fifth of a penny each, before any detailed reading was spent on them.

Where did it go wrong?

This is the part most benchmarks leave out.

Till receipt line items: 78.9%. On photographed thermal receipts, about four lines in five were read with every column right. The header and the total on the same receipts were read perfectly, so the amount that reaches your records was correct; it is the line-by-line breakdown that suffers. Thermal print is small, low in contrast and often crumpled, and that shows.

Scanned PDF line items: 84.8%. Just under the 85% bar for that format on its own, though the corpus as a whole clears it. Core fields on the same scans were 100%.

Ten invented values across 116 documents. A VAT rate was returned seven times on documents that print none, and an invoice number three times on documents that print none. We count those against the pipeline. In the app, each field carries a confidence and anything doubtful is put in front of the user to check, which is what catches this kind of error before it reaches a tax figure.

Two extra line items were returned that the pages do not print.

“The totals were right on every till receipt. The line-by-line breakdown was right four times in five. Both of those are true, and you need both to judge a receipt scanner.”
TapTax, Benchmark note

How did each kind of document do?

FormatCore fieldsLine itemsMean cost
Text PDF100.0%100.0%1.39p
Multi-page text PDF100.0%100.0%3.10p
Handwritten100.0%100.0%3.53p
Multi-page scanned PDF100.0%100.0%11.38p
Scanned PDF100.0%84.8%4.27p
Photo99.2%94.4%2.68p
Till receipt photo100.0%78.9%2.78p

The pattern is the one you would expect. A PDF straight from accounting software reads perfectly and cheaply, because the text is already there. Images cost more to read and lose accuracy on fine detail. The handwritten receipts did better than we expected, though six documents is too few to draw a firm conclusion.

By document type, invoices scored 99.8% on core fields and 97.6% on line items, receipts 100% and 81.0%, and credit notes 100% on both.

How does the reading actually work?

Each document is first classified from its first page or a small preview: invoice, receipt, credit note, or not financial. A PDF with a text layer is then read by a fast, low-cost model; an image or scan goes to a stronger vision model; and any field that comes back doubtful can be escalated to the most capable model. The models are Anthropic's Claude family, which our privacy policy names as the processor.

After reading, every figure is checked by ordinary code, not by the model: net plus VAT must equal the total, the VAT rate must be a UK rate, and the dates must make sense. A document that fails a check is flagged rather than corrected by guesswork. A transient failure is retried up to three times, as the production job system does, and every retry is included in the cost.

What are the limits of this benchmark?

  • The documents are synthetic. That protects real people's data, but real receipts are messier than any generator: torn, folded, faded, photographed in a van at dusk. Expect real-world accuracy to be lower, most of all for photographed receipts.
  • 116 documents shows the shape, not a margin of error. Groups as small as six handwritten receipts or two multi-page scans are indicative only.
  • It measures reading, not judgement. Whether a cost is allowable, or which HMRC expense category it belongs in, is not scored here. That is a separate question, covered in our guide to what expenses you can claim.
  • It measures our pipeline. Other products may do better or worse; we have not tested them, and vendor accuracy claims we quote elsewhere are theirs.

Why publish this at all?

Under Making Tax Digital, sole traders and landlords with qualifying income over £50,000 keep digital records from 6 April 2026, with the threshold falling to £30,000 in April 2027 and £20,000 in April 2028. For many of them, the receipt scanner is now part of how their tax return is built. A number on a marketing page is not enough to trust that with; a method, a corpus and the failures are.

We intend to re-run the benchmark as the pipeline changes and publish each run, including any that go backwards.

Try it on your own documents

You can run the same reader on your own receipts and invoices, without an account, and see each field and its check:

Neither keeps your file. In TapTax itself, receipts and bills are read, checked, matched to your bank statement and filed with your quarterly update; the receipts and bills page shows how.

People also ask

How accurate is AI at reading receipts?

In our September 2026 benchmark of 116 synthetic UK documents, the reader got 99.8% of core fields (supplier, date, net, VAT, total) and 95.6% of line items right. Photographed till receipts were the weak spot, at 78.9% for line items, though their totals were all correct.

Is a scanned receipt good enough for HMRC?

HMRC accepts clear digital copies of records. A scanned receipt that shows the supplier, date, amount and any VAT is a valid record, and you must keep business records for at least 5 years after the 31 January deadline for the tax year.

Why test on synthetic documents rather than real ones?

So the whole benchmark can be published without exposing anyone’s financial records. The trade-off is that real documents are messier, so real-world accuracy will be lower, especially for photographed receipts.

What does it cost to read a receipt with AI?

In this run, 2.61p a document on average at list API prices, from 1.39p for a digital PDF to 11.38p for a two-page scan. Documents that turned out not to be financial cost about 0.2p to recognise and set aside.

Topicsreceipt scanninginvoice OCRAI accuracyexpense receiptsMaking Tax DigitalTapTax research
ShareXLinkedIn

Written by

Solomon Amos

Founder, TapTax

Solomon is a tax technology expert and the founder of TapTax. He writes plain-English guides on Making Tax Digital, HMRC compliance, and UK sole trader taxes - because everyone deserves to understand their own tax obligations.

  • Built TapTax's HMRC Making Tax Digital integration
  • Three years as a lead technical architect on HMRC digital programmes
  • PhD in engineering

Published

More about Solomon
  • MTD Guides

    Automatic Receipt Scanning Tax UK: Does It Actually Work?

    Automatic receipt scanning for UK tax sounds like a dream. Here's what it actually does, what it misses, and whether it's worth it for sole traders.

    8 min read

  • MTD Guides

    Invoice Finance and MTD: The Record-Keeping Trap

    Using invoice finance as a sole trader creates a hidden MTD record-keeping problem. Here is what HMRC expects and how to avoid a costly compliance gap.

    8 min read

  • Tax Tips

    Invoice Finance for Sole Traders: The Fee Nobody Quotes

    An invoice finance facility charges 1-5% per invoice to unlock your cash early. Here is what UK sole traders actually pay, and why your records set the rate.

    8 min read

Free calculators for the figures this article talks about.

Stop dreading your tax return.

TapTax imports your bank statements, categorises expenses automatically, and submits quarterly updates to HMRC. Free plan, no card required.

Get started free