How AI Recognizes Invoices, Acts, and Bills: OCR + LLM in Simple Terms
How modern document recognition differs from old template-based OCR, how the system checks amounts and company ID, and why a human in the loop is still necessary.
In short, if you don’t have time to read it all
- ✓Old OCR systems relied on templates and broke with every new format.
- ✓Multimodal models understand the structure of the document, not just the coordinates of the fields.
- ✓Each field is checked against rules: control number of the company ID, IBAN format, arithmetic of amounts.
- ✓If the system is unsure, the document goes to the accountant. Payments cannot be processed without a human.
The accountant opens their inbox. There’s an invoice in PDF, a photo of a waybill taken on a phone in the back of a truck, and a scanned acceptance certificate with the stamp landing right on top of the amount. And all of it has to be entered by hand into the accounting system. Every single day.
People have been trying to automate this for a long time. But for years, the results were pretty underwhelming.
1. Why old OCR never really took off
Classic recognition systems worked off templates. For each supplier, you had to configure them: amount — here, document number — here, date — top right corner. The supplier changed the form — the template broke. A new counterparty showed up — time to sit down and configure it again.
For ten regular suppliers, that’s still manageable. For hundreds, it isn’t.
2. What language models changed
Modern multimodal models look at a document roughly the way a person does. They understand that “Total due” is the final amount, no matter where it appears. That a table with columns like “Item”, “Quantity”, and “Price” is the line-item list. That VAT (locally called PDV) can be written in a dozen different ways.
But that alone still isn’t a solution. The model can make mistakes. So recognition is only the first step.
3. How the system checks itself
After the model extracts the fields, each one goes through rule-based validation. Plain, ordinary checks — no AI involved:
- Company ID — the check digit is validated. A one-digit mistake gets caught immediately.
- IBAN — the format and checksum are validated.
- Arithmetic — the line totals must add up to the grand total, and the amount before VAT plus VAT must equal the overall total.
- Reference data — the supplier is looked up in your counterparty database, and the item list is matched against your catalog.
- Contract prices — if the price differs from the one fixed in the specification, the document is flagged.
4. Confidence threshold and Human-in-the-loop
The system assigns a confidence score to every field. If even one field falls below the threshold, or the amount exceeds the set limit, or the checks don’t match up — the document does not go through automatically. It gets highlighted for the accountant, with the recognized fields shown right on top of the scan. Checking and correcting it takes just a few seconds.
And most importantly: the system prepares a draft. Money is never paid out from an invoice without sign-off from the responsible person. Never.
5. What you need to get started
- A sample of real documents from your key counterparties — a few dozen is enough.
- Access to your accounting system or document workflow system via API (for example, DocuSign or QuickBooks).
- Counterparty and item reference data.
- Rules: what should be considered suspicious, and who confirms it.
See how it works step by step in the document recognition and approval demo scenario. More about the service is on the Document Processing Automation page. And if you want to estimate how much time this could free up for your accounting team, try the calculator.
- Procedure for Calculating the Control Number of the Company ID
- IBAN Standard (ISO 13616)
Document Processing Automation
Still have questions about the article topic?
Let’s look at how these approaches fit your company’s actual processes.