About

Built because generic OCR kept getting Arabic and Hebrew wrong

RTLDocs started in 2024 as an internal tool. We were building document pipelines for fintech and legal clients and kept hitting the same wall: OCR engines built LTR-first would silently reverse digits, mirror punctuation, and scramble table columns on Arabic and Hebrew documents. The JSON always looked fine — until someone actually read it.

2024
Founded
40k+
Documents benchmarked
94.2%
Arabic accuracy
91.8%
Hebrew accuracy

Why we started

Patching a general-purpose OCR engine with regexes and manual review queues doesn't scale, and it doesn't fix the root cause: most extraction pipelines never model bidi text, letter joining, or script-aware digit forms in the first place. We built a dedicated RTL-aware pipeline instead of another workaround.

What we built

Layout reconstruction, bidi correction, and script-aware OCR run together on every document, in one pass. The output is typed, validated JSON with a confidence score and a direction tag on every text run — so RTL fields render correctly the first time, in your app and in ours.

Who we work with

Fintech teams verifying Arabic and Hebrew ID documents at onboarding speed, ERP and accounting systems ingesting invoices and receipts in any script, legal teams extracting structured terms from bilingual contracts, and AI teams grounding agents and RAG pipelines in document truth.

Where we're headed

Arabic and Hebrew are first-class today; the same pipeline extends naturally to the rest of the RTL script family. We're also expanding self-hosted deployment options for teams with data residency requirements.