Flovvy

Why we believe in what we built

Internal test · run September 2026 · app version 1.0.0

This is not a promise — it is the test we ran on ourselves. Detailed receipt scanning took a long time to get right, and it works by combining several techniques rather than leaning on one. We will keep improving it. But before putting it in anyone’s hands we wanted to know whether the thing we were about to sell actually worked — so we built a bench, transcribed the truth by hand, and counted. This page reports what came out: the results, how the bench was made, and where it falls short.

What the test showed

ResultWhat was measuredOut of
281 of 281The receipt total, read correctly281 readings
99.9%Product lines with the right amount3,204 of 3,207 lines
98.9%The product lines add up to that total278 of 281 readings
95.1%Right amount and right description3,049 of 3,207 lines
87 of 90Weighed items, read correctly90 weighed lines
444Extra pack formats read out of the description itself — “GR 500”, “LT.1”on top of the 90 weighed lines
461 of 462Multiples we reported, read correctly462 reported multiples
zeroSingle items turned into multiples, on the lines the receipt shows as one piece2,159 lines

The most delicate part is the product descriptions. On a poor photograph the wording is what degrades first, and often the difference is a single letter.

This looks like a strong result to us, and a solid enough basis to be comfortable shipping the feature — which is what the test was for.

How the bench was built

It is 254 receipts — 281 readings, because some were photographed more than once, on purpose, to see whether a different shot changes the answer — covering 3,207 product lines, from over 25 countries and 25 currencies.

It was assembled to be difficult, not representative of a good day. It mixes supermarket tickets with restaurant and bar bills, clothing shops, toy shops and other retail; receipts of one line and receipts of eighty; line discounts, multi-buys, weighed items and two VAT rates on the same document. 66 of the 281 readings are receipts printed in a non-Latin script — Greek 38, Chinese and Japanese 14, Cyrillic 12, kana 8, Arabic 4, Hangul 3.

The photographs are awkward on purpose. Many are shot at an angle or rotated, creased, badly lit or plainly low quality — the way a receipt actually gets photographed at the end of a shop, not the way it gets photographed for a demo. The only ones left out were those a person cannot read either: if the printing had faded past human legibility, scoring a machine against it measures nothing. Nothing was dropped for being hard, and the state of the photographs is measured below rather than asserted.

What the photos are actually like

Since the state of the photograph is where a good part of that comes from, we measured it rather than describing it, image by image:

FigureWhat it measures
86 / 277 · 31%Photographs taken with the phone rotated — roughly one in three, across the 277 image files in the bench.
59 / 229 · 26%Receipts sitting crooked in the frame by more than a degree, up to 29°. Measured on the 229 images where the paper could be isolated from the background.
53%Of the images whose outline could be measured, the share of receipts that do not lie flat in their own rectangle — rolled, creased or folded. Half of those are badly out of shape.
56 / 281Readings whose hand transcription explicitly notes a defect in the photograph: cut off or out of frame 25, rotated or skewed 14, torn, stained or faded 8, blurred 5, folded 5, background or light 4. A lower bound — the note only gets written when the defect caused trouble.
53 lines on 26 receiptsProduct lines where the printed wording simply is not legible in the photograph. No rule can guess those, and we do not pretend to.

One photograph per receipt. Every reading in the bench is a single shot of a whole receipt — never a long receipt split in two and photographed in pieces, or several frames stitched side by side. That is deliberate: the tool is built for the way a receipt is actually photographed, in one go, and we do not intend to change that.

What this does not say

A few things, stated plainly because a page like this is worth nothing if it only reports the good half.

Why this page has a date

Because a test does. These figures belong to one run of the bench and one version of the app, both stated at the top — they are not a permanent property of the product. We will keep working on the reading engine, and on the bench that measures it, to make the service better than what this page reports.