Benchmarked on 40k+ RTL-language documents

Document extraction that actually works on RTL languages.

Structured JSON from invoices, contracts, and statements — in 100+ languages, with special attention paid to right-to-left scripts. One API for your entire document flow, benchmarked against Google and Azure.

95%+
RTL accuracy
100+
Languages supported
40k+
Documents benchmarked
extract.sh
curl -X POST https://api.rtldocs.ai/v1/extract \
-H "Authorization: Bearer sk_live_***" \
-F "file=@invoice.pdf"
{
"fields": {
"vendor_name": "Global Trading Co.",
"invoice_number": "INV-2026-0142",
"total": "1,587.00",
"currency": "USD"
},
"confidence_summary": { "mean": 0.98 }
}
Playground

Try it on a real document

Drop a file, or pick a sample below — invoice, contract, or statement, in any language. No signup required.

Drop a file here, or click to browse
PDF, PNG, JPG — up to 20MB
Or try a sample document
The problem

Why generic OCR fails on RTL

Most extraction engines are built left-to-right first. Bidirectional text, mirrored punctuation, and connected scripts break silently: the JSON looks valid until someone actually reads it.

Reversed digits and numbers

Numbers, dates, and amounts inside right-to-left sentences come out reordered or fully reversed.

Mirrored punctuation

Parentheses and brackets flip direction, and quotation marks land on the wrong side of the text they wrap.

Broken letter joining

Connected scripts like Arabic fragment into isolated letterforms, producing text that can't be read, searched, or matched.

Scrambled table columns

Column order flips in bilingual tables, invoices, and bank statements, so values end up under the wrong headers.

Garbled names and identifiers

Latin names, IDs, and reference codes embedded in RTL text come out split apart or in the wrong order.

Reordered lines and paragraphs

It's not just characters. Whole lines and paragraphs can come out in the wrong sequence.

How it works

Three steps, one endpoint

POST your document
Send a PDF or image to a single endpoint — any language, any layout.
RTL-aware extraction pipeline
Layout reconstruction, bidi correction, and script-aware OCR run together.
Typed, structured JSON
Get validated fields with confidence scores and direction metadata.
Your appPDF / imageExtraction API/v1/extractRTL pipelineOCR · bidi · layoutStructured JSONfields + confidence
Use cases

Built for document-heavy workflows

Fintech
KYC & onboarding
Verify RTL-language ID documents at fintech onboarding speed.
ERP / accounting
Invoices & receipts
Feed ERP and accounting systems structured line items, in any script.
Legal tech
Contracts & legal
Extract parties, dates, and clauses from bilingual legal documents.
AI infrastructure
Agent & RAG pipelines
Ground agents in document truth with typed fields. MCP server available.
Pricing

Simple, usage-based pricing

MonthlyAnnual (2 months free)
Free
$0
100 pages / month
  • ✓Playground access
  • ✓Full REST API
  • ✓Community support
Start free
Most popular
Growth
$49/mo
2,000 pages / month, $0.02/page overage
  • ✓Everything in Free
  • ✓Priority queue
  • ✓Webhook callbacks
Get started
Scale
$299/mo
20,000 pages / month
  • ✓Everything in Growth
  • ✓Priority support
  • ✓Self-hosted option — contact us
Get started
Integrate

Same call, every stack

curl -X POST https://api.rtldocs.ai/v1/extract \
  -H "Authorization: Bearer sk_live_***" \
  -F "file=@invoice.pdf"
FAQ

Common questions

Documents are processed in-memory and deleted immediately after extraction. Zero-retention is the default on every plan — nothing is stored or used for training unless you opt in.
Contact us

Contact us

Questions about pricing, self-hosting, or whether your document type is a good fit? Send us a note — we read every message ourselves.