What AI gets wrong on an Israeli receipt
Photograph a receipt, get a filled expense form back. The feature works, and on a printed A4 invoice from a large supplier it is close to perfect. The failures are not random, though. They cluster in six places, and five of the six are specific to how receipts look in Israel.
Hebrew on thermal paper
Character recognition is strongest on printed Latin text and weakest on exactly what a small Israeli supplier hands over: Hebrew, low contrast, a dot matrix head that skipped half a line, ink already fading. The amount usually survives because digits are easy. The supplier name is what comes back mangled, and a wrong supplier name is what makes the same vendor split into four entries in your books. Check the name field, not the total.
A VAT line that was calculated, not read
Plenty of small receipts print one number and no VAT breakdown. Faced with that, a model does the arithmetic at 18 percent and returns a confident VAT amount that appears nowhere on the paper. If the supplier is an exempt dealer there is no VAT on the transaction at all, and the number is invented. Nothing on the image says which case you are in, so the model cannot know either.
03/04/2026
Third of April or fourth of March. Israeli receipts are day first, a model trained mostly on American documents leans month first, and the two readings fall into different bi-monthly VAT periods. The error is invisible until a period closes with a total that does not match.
A card slip is not a tax invoice
Both are small pieces of paper with an amount, and a model will happily describe either as a receipt. Only the tax invoice supports an input VAT deduction. A slip is evidence that money moved. Anything that reads the two the same way will build you a VAT claim with no document behind part of it.
The deductible percentage is law, not text
A category is a reading task. The percentage attached to it is not. How much of a meal, a car or a phone bill is deductible comes from legislation that changes, and it belongs in a maintained table, looked up after the category is chosen. A model asked to produce the percentage will produce a plausible one.
Confidence is not accuracy
Scores measure how clean the input looked. A crisp receipt read wrongly scores high. Treat a low score as a reason to look and a high score as no reason to skip looking.
The check that catches all six
Every one of these is caught by the same habit: the model fills the form, you read the four fields that matter, then you save. Supplier, date, total, VAT. Ten seconds, and it keeps the speed while removing the part that quietly corrupts a year of records. Software that saves a scan without showing you those fields first is optimizing its own demo.
Will a better model fix this
For two of the six, yes. Hebrew on faded thermal paper is an image quality problem, and image quality problems get better every year. Ambiguous dates are partly solvable too, because a system that knows the document is Israeli can apply the local convention instead of guessing.
The other four are not model problems and will not be fixed by a larger one. A VAT line that does not exist on the paper cannot be read off it, at any level of capability. Whether a slip is a tax invoice is a legal classification, not a visual one. The deductible percentage lives in legislation. Confidence measures the wrong thing by construction. Vendors who describe all six as accuracy issues awaiting the next model release are describing two of them correctly.
Where errors hide: batch scanning
Scanning one receipt at the counter is the safe case, because you are holding the paper and the screen at the same time. The dangerous flow is the end-of-month batch, twenty photographs uploaded together and confirmed with one button. Nobody reads twenty forms carefully. The interface makes it feel checked, and the errors go in with an approval attached.
If you work in batches, sort the review by amount and read the large ones properly. An error on a 38 shekel parking receipt costs you almost nothing. The same error rate applied to a 9,000 shekel equipment invoice is the one that matters, and it is one row out of twenty.
What a well-built review screen looks like
- The image and the extracted fields sit side by side, so checking does not mean remembering.
- Fields the model was unsure about are marked, and the marking is on the field rather than one score for the document.
- A VAT amount that was calculated rather than read is labelled as such, or is left empty.
- Nothing saves until a person presses save, and the button does not say confirm all.
Every one of those is a design decision, not a technical limitation. Software that gets them wrong is optimizing the demo, where speed is visible and a wrong supplier name is not.
The supplier name problem, specifically
Of the six, the one that quietly does the most damage is the mangled supplier name, because nothing about it looks wrong. The amount is right, the date is right, the expense is in the books. What happens is that the same petrol station arrives as four slightly different names across a year, and the fuel total you look at in December is spread across four rows that never get added together.
The fix is boring and effective: keep a supplier list, and match new scans against it rather than creating a new supplier every time. Any system that lets a scan invent a supplier record with no matching step will accumulate this problem invisibly, and you will only find it when a year of data refuses to summarize.
This is general information, not tax advice for your situation. Rates and rules change, so confirm current figures before acting. Next: what AI really does in accounting, or try the expense categorizer.