Internal test · run September 2026 · app version 1.0.0
| Result | What was measured | Out of |
|---|---|---|
| 281 of 281 | The receipt total, read correctly | 281 readings |
| 99.9% | Product lines with the right amount | 3,204 of 3,207 lines |
| 98.9% | The product lines add up to that total | 278 of 281 readings |
| 95.1% | Right amount and right description | 3,049 of 3,207 lines |
| 87 of 90 | Weighed items, read correctly | 90 weighed lines |
| 444 | Extra pack formats read out of the description itself — “GR 500”, “LT.1” | on top of the 90 weighed lines |
| 461 of 462 | Multiples we reported, read correctly | 462 reported multiples |
| zero | Single items turned into multiples, on the lines the receipt shows as one piece | 2,159 lines |
The most delicate part is the product descriptions. On a poor photograph the wording is what degrades first, and often the difference is a single letter.
This looks like a strong result to us, and a solid enough basis to be comfortable shipping the feature — which is what the test was for.
It is 254 receipts — 281 readings, because some were photographed more than once, on purpose, to see whether a different shot changes the answer — covering 3,207 product lines, from over 25 countries and 25 currencies.
It was assembled to be difficult, not representative of a good day. It mixes supermarket tickets with restaurant and bar bills, clothing shops, toy shops and other retail; receipts of one line and receipts of eighty; line discounts, multi-buys, weighed items and two VAT rates on the same document. 66 of the 281 readings are receipts printed in a non-Latin script — Greek 38, Chinese and Japanese 14, Cyrillic 12, kana 8, Arabic 4, Hangul 3.
The photographs are awkward on purpose. Many are shot at an angle or rotated, creased, badly lit or plainly low quality — the way a receipt actually gets photographed at the end of a shop, not the way it gets photographed for a demo. The only ones left out were those a person cannot read either: if the printing had faded past human legibility, scoring a machine against it measures nothing. Nothing was dropped for being hard, and the state of the photographs is measured below rather than asserted.
Since the state of the photograph is where a good part of that comes from, we measured it rather than describing it, image by image:
| Figure | What it measures |
|---|---|
| 86 / 277 · 31% | Photographs taken with the phone rotated — roughly one in three, across the 277 image files in the bench. |
| 59 / 229 · 26% | Receipts sitting crooked in the frame by more than a degree, up to 29°. Measured on the 229 images where the paper could be isolated from the background. |
| 53% | Of the images whose outline could be measured, the share of receipts that do not lie flat in their own rectangle — rolled, creased or folded. Half of those are badly out of shape. |
| 56 / 281 | Readings whose hand transcription explicitly notes a defect in the photograph: cut off or out of frame 25, rotated or skewed 14, torn, stained or faded 8, blurred 5, folded 5, background or light 4. A lower bound — the note only gets written when the defect caused trouble. |
| 53 lines on 26 receipts | Product lines where the printed wording simply is not legible in the photograph. No rule can guess those, and we do not pretend to. |
One photograph per receipt. Every reading in the bench is a single shot of a whole receipt — never a long receipt split in two and photographed in pieces, or several frames stitched side by side. That is deliberate: the tool is built for the way a receipt is actually photographed, in one go, and we do not intend to change that.
A few things, stated plainly because a page like this is worth nothing if it only reports the good half.
Because a test does. These figures belong to one run of the bench and one version of the app, both stated at the top — they are not a permanent property of the product. We will keep working on the reading engine, and on the bench that measures it, to make the service better than what this page reports.