Document extraction that actually works on RTL languages.
Structured JSON from invoices, contracts, and statements — in 100+ languages, with special attention paid to right-to-left scripts. One API for your entire document flow, benchmarked against Google and Azure.
Try it on a real document
Drop a file, or pick a sample below — invoice, contract, or statement, in any language. No signup required.
Why generic OCR fails on RTL
Most extraction engines are built left-to-right first. Bidirectional text, mirrored punctuation, and connected scripts break silently: the JSON looks valid until someone actually reads it.
Numbers, dates, and amounts inside right-to-left sentences come out reordered or fully reversed.
Parentheses and brackets flip direction, and quotation marks land on the wrong side of the text they wrap.
Connected scripts like Arabic fragment into isolated letterforms, producing text that can't be read, searched, or matched.
Column order flips in bilingual tables, invoices, and bank statements, so values end up under the wrong headers.
Latin names, IDs, and reference codes embedded in RTL text come out split apart or in the wrong order.
It's not just characters. Whole lines and paragraphs can come out in the wrong sequence.
Three steps, one endpoint
Built for document-heavy workflows
Simple, usage-based pricing
- ✓Everything in Free
- ✓Priority queue
- ✓Webhook callbacks
- ✓Everything in Growth
- ✓Priority support
- ✓Self-hosted option — contact us
Same call, every stack
curl -X POST https://api.rtldocs.ai/v1/extract \ -H "Authorization: Bearer sk_live_***" \ -F "file=@invoice.pdf"
Common questions
Contact us
Questions about pricing, self-hosting, or whether your document type is a good fit? Send us a note — we read every message ourselves.