OCR + LLM Document Classification: Saudi Finance
How we built a document pipeline that classified and parsed 12,000+ Saudi financial documents, going from 30% accuracy with OCR alone to a 98.5% median confidence with OCR plus vision models.
Every accounting firm has the same problem, and it isn't accounting. It's finding things. Client documents live in email threads, shared drives, WhatsApp chats and phone photos of printed invoices. When a filing deadline or an audit comes around, accountants spend the first days hunting for paper instead of reading it.
Nemra was our answer for Saudi Arabia: a platform that collects financial documents from wherever they live, works out what each one is, extracts the data, and turns it into structured, checkable records an accounting system can use. Before Nemra was acquired, it had processed more than 12,000 Saudi financial documents. This is how the product worked, and the engineering decisions that got it there.
Why Saudi Arabia, and why one vertical
Saudi Arabia was a deliberate bet. Under Vision 2030 the country is digitizing its own compliance infrastructure, and ZATCA, the tax authority, was pushing e-invoicing hard. Its fintech sandbox was open to builders, and regulators were reachable. A compliance-heavy product had a market that was moving toward it.
We went narrow on purpose: accounting firms and in-house finance teams with high document volumes. Not "AI for finance", one job done properly. That focus shaped every technical decision below, and it's what let a small team ship: five people in Algeria and two on the ground in Saudi Arabia.
Staying close to the regulator turned out to be an advantage of its own. Our checks followed ZATCA's rules as they changed, not months later, and clients noticed.
The product: collect, understand, check
Nemra had three layers, each isolated from the others:
- Extraction. Every document, whatever its format, becomes structured data: document type, parties, amounts, tax, line items, dates.
- Audit. Rule-based and model-assisted checks run on that data and flag inconsistencies, missing fields and anomalies before an accountant has to look.
- Storage. Financial data is sensitive, so storage was built on the assumption that what goes in stays in, with a full audit trail on every record.
The interesting engineering, and most of the work, was in the first layer.
Why OCR alone failed: 30% accuracy
My first version did what everyone does first: extract the text from each PDF with PyMuPDF, run it through a text classifier, and pull fields out with regex templates. On the first test documents it worked perfectly.
Then we tested on real documents from real Saudi businesses, and classification accuracy fell to around 30%. The reasons were consistent:
- Every document is a variant. Each bank has its own statement layout and each supplier designs its own invoice. Template rules that worked for one format broke on the next one that was only slightly different, and adding templates became whack-a-mole.
- OCR output is messy. Low-quality scans, handwritten notes, stamps, and multi-column layouts produced text that was garbled or out of order. A classifier trained on clean text didn't know what to do with it.
- Text throws away what a human sees first. A logo, a coloured header, a table's position on the page: that's how a person tells a payslip from an invoice at a glance. Plain text loses all of it.
- Too many moving parts. OCR services, cleaning scripts, rate limiting and error handling scattered across the pipeline meant every connection point could fail, and I spent more time debugging plumbing than improving results.
Training a custom image classifier would have needed thousands of labelled documents and the infrastructure to train on them. Multimodal language models offered a shortcut: models that already read documents, if you give them the right inputs.
The pipeline that worked: text and image together
The redesign rests on one rule: treat every document as both text and image, and let a vision-capable model combine the two.
Ingestion. Documents arrive through a FastAPI endpoint, as PDFs or images. For classification, the first two pages almost always carry the deciding clues: headers, logos, the document's title.
A hybrid representation. The first pages are rendered to high-resolution images and run through OCR, keeping the bounding boxes so the layout survives. The page images are stacked vertically into a single image, encoded in base64, and sent together with the extracted text. The model sees that something looks like an invoice and can read the text to confirm it.
Classification with an exit. The model sorts each document into a fixed set of types (bank statement, invoice, receipt, payslip, cheque, transfer note, other) and returns a confidence score. When it isn't sure, it must answer "other" rather than guess, and those documents go to a human.
pythonasync def classify_document(page_image_b64: str, ocr_text: str) -> dict: response = await llm.chat.completions.create( model=settings.LLM_MODEL, response_format={"type": "json_object"}, messages=[{ "role": "user", "content": [ {"type": "text", "text": CLASSIFY_PROMPT.format(ocr_text=ocr_text[:2000])}, {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{page_image_b64}"}}, ], }], ) return ClassificationResult.model_validate_json(response.choices[0].message.content)
Long documents. Multi-page statements are split into chunks of about five pages. Each chunk is summarized on its own, and a second call combines the summaries, which keeps every request within the model's limits without losing data buried on page nine.
Extraction by type. Once a document is classified, a dedicated parser takes over, with its own prompt and a Pydantic schema the output must satisfy. The invoice parser extracts the supplier, invoice number, line items, amounts and VAT, then produces the double-entry journal entries against the client's chart of accounts, which is the part accountants actually wanted. If the model returns anything that doesn't fit the schema, it's caught immediately instead of reaching the books.
The result: 98.5% median confidence
Across every document the platform processed, the median classification confidence reached 98.5%, up from roughly 30% accuracy with OCR alone. Just as important, the uncertain cases announced themselves: low-confidence documents were routed to review instead of being filed wrong. Layout changes stopped breaking things, because nothing depended on templates any more, and supporting a new document type meant adding a prompt and a schema, not retraining a model.
Provider-agnostic by design
In production the pipeline ran on GPT-4o, but nothing in it depended on OpenAI. Every model call went through an OpenAI-compatible client whose base URL and model name came from configuration, so switching to another provider, or to a self-hosted model behind a compatible API, was a configuration change.
That mattered for three reasons:
- Cost and quality move fast. Being able to swap models meant we could take a better or cheaper one the week it shipped.
- Data residency. Financial data and regulated markets raise questions about where processing happens. A compatible endpoint inside the right jurisdiction keeps the whole pipeline unchanged.
- Testing. Mock and staging endpoints plug into the same client, so the pipeline could be exercised end to end without spending tokens.
ZATCA in practice
The pipeline earned its keep on e-invoicing. For one client, a construction contractor on heavily customized invoicing software, Nemra became the compliance layer around their existing system instead of a replacement for it: they exported invoices as PDFs, Nemra extracted and validated them against ZATCA's rules, generated the structured XML with the required signature and stamp, and submitted them to ZATCA's Fatoora portal. I wrote that project up separately in ZATCA Phase 2 middleware for legacy systems.
What I took from building Nemra
- Go deep on one vertical. A product that does one job better than anyone else is easier to build, easier to sell and easier to trust than one that does many jobs adequately.
- OCR is one signal, not the answer. Combining what a document says with how it looks beat every amount of template tuning.
- Make uncertainty a feature. A system that says "I'm not sure" and routes the document to a person is worth more than one that is confidently wrong.
- Validate everything with schemas. Model output is only useful once it's been forced into a shape your accounting system can check.
- Stay close to the regulator. In compliance products, being current is the product.
The same philosophy now powers ComptaLegal for Algerian fiscal compliance: different regulations, same idea. Meet companies where they already work, and put the intelligence in the layer around their tools.