API reference

RTLDocs API

RTLDocs exposes a small, typed extraction API. Send a document, get back structured JSON — every text run tagged with a script, a reading direction, and a confidence score, correctly ordered even for Arabic and Hebrew content embedded inside an otherwise left-to-right response.

Quickstart

POST a file as multipart form data to /v1/extract.

curl -X POST https://api.rtldocs.ai/v1/extract \
  -H "Authorization: Bearer sk_live_***" \
  -F "file=@invoice_ar.pdf"

Extracted values live under fields, as plain strings. The confidence score and direction tag live one level deeper, on each text run in pages[].runs[] — RTL values render correctly even inside this LTR JSON block.

{
"fields": {
"vendor_name": "شركة النور للتجارة",
"total": "١٬٥٨٧٫٠٠"
},
"pages": [{ "runs": [{
"text": "شركة النور للتجارة",
"direction": "rtl",
"confidence": 0.97
}] }]
}

Authentication

Every request needs a bearer token — an API key generated from your dashboard. Send it as an Authorization header on every request.

Authorization: Bearer sk_live_***

Missing or invalid keys get a 401 with an ErrorResponse body. Keys are scoped to your account's plan and monthly page quota.

Extract a document

POST /v1/extract

POST/v1/extractExtract structured fields from a document

Accepts either multipart/form-data (upload a file, or pass a url field instead) or application/json (fetch a document from a public url — no upload). The two content types share the same option fields below; the JSON body only differs in that languages is an array of strings there instead of a comma-separated string, and url is required.

Request parameters

FieldTypeDescription
filebinaryThe document to extract (PDF, PNG, or JPEG). multipart/form-data only — mutually exclusive with url.
urlstringPublicly fetchable URL of the document. Required when Content-Type is application/json; optional alternative to file when uploading multipart.
document_typeenumOne of auto, invoice, receipt, contract, bank_statement, id_document, generic. auto detects the type for you.Default: "auto"
languagesstring[]Hint the expected languages, e.g. ar,en (multipart form field) or ["ar","en"] (JSON array). Optional — omit to let the model detect languages per run.
outputenumfull returns fields, tables, per-page text runs and the full logical-order text. fields_only returns just the fields object. text_only returns just full_text_logical_order.Default: "full"
modeenumsync blocks until extraction finishes and returns the result directly (200). async returns 202 immediately with a job_id — poll GET /v1/jobs/{job_id} for the result.Default: "sync"
formatenumjson returns the full ExtractResponse object. txt/md return the extracted text as a raw text/plain or text/markdown body instead — sync only. For mode=async, the 202 is always JSON; pass format as a query param on the job-status endpoint instead.Default: "json"

Response — ExtractResponse (200, sync)

Returned when mode=sync and format=json (the default). If format is txt or md, the body is the extracted text as a raw string instead.

FieldTypeDescription
request_id*stringUnique identifier for this extraction call.
document_type*stringThe resolved document type (echoes your request, or the auto-detected type).
pagesPageResult[]Per-page breakdown — text runs, tables, and warnings for each page. See Pages & text runs.
fieldsobjectThe extracted key/value fields, typed per document (e.g. vendor, total, invoice_number). Shape varies by document_type.
additional_fieldsRecord<string,string>Any extra fields the model found that don't map to a known field name, as raw strings.
tablesTable[]Document-level tables (line items, statement rows, etc), each a grid of cells with row/column spans.
full_text_logical_orderstringThe complete extracted text in correct reading order — bidi-corrected, so RTL and LTR spans read naturally even when mixed.Default: ""
confidence_summary*ConfidenceSummaryAggregate confidence stats across every run. See Confidence summary.
warningsstring[]Document-level warnings (e.g. low-quality scan, partially unreadable page).
processing_ms*integerWall-clock time the extraction took, in milliseconds.
model*stringThe model that processed this document, e.g. gemini-flash-latest.
prompt_version*stringVersion tag of the extraction prompt used — useful for reproducing results.

PageResult & ProcessedTextRun

Each entry in pages is a PageResult:

FieldTypeDescription
page_number*integer1-indexed page number.
widthinteger | nullPage width in source units, if known.
heightinteger | nullPage height in source units, if known.
runsProcessedTextRun[]Every text run detected on this page, in reading order.
tablesTable[]Tables detected on this page.
warningsstring[]Page-level warnings.

Each entry in runs is a ProcessedTextRun — a span of text after script-aware postprocessing (bidi fixes, Unicode normalization, digit policy):

FieldTypeDescription
text*stringThe run's text, ready to display.
direction*enumltr or rtl — this run's natural reading direction.
language*stringISO 639-1 code (e.g. ar, he, en), or "und" if undetermined.
script*enumOne of arabic, hebrew, latin, cyrillic, greek, other.
confidence*number (0–1)Model confidence for this run.
bboxBoundingBox | nullPixel bounding box on the page — x0, y0, x1, y1 — if the source provided positional data.
raw_text*stringText exactly as printed/extracted, before any cleanup.
normalized_text*stringNFC-normalized, bidi-corrected, logical-order text.
digit_formenumNumeral system used: arabic_indic (١٢٣), western (123), mixed, or none.Default: "none"
flagsstring[]Non-destructive issue markers on this run (e.g. low confidence, possible OCR ambiguity).

ConfidenceSummary

FieldTypeDescription
meannumberMean confidence across every run in the document.Default: 0.0
minnumberLowest confidence score among all runs.Default: 0.0
low_confidence_run_countintegerNumber of runs below the low-confidence threshold — worth a human review pass.Default: 0
total_run_countintegerTotal number of text runs in the document.Default: 0
Async extraction

GET /v1/jobs/{job_id}

GET/v1/jobs/{job_id}Check status of an async job

Poll this after a POST /v1/extract with mode=async returned a job_id.

FieldTypeDescription
job_id*string (path)The job ID returned by the 202 response.
formatenum (query)json returns the full JobStatusResponse. txt/md return just the extracted text as a raw body instead — only once status is completed. While pending, processing, or failed, the JSON status is always returned regardless of this param.Default: "json"

Response — JobStatusResponse:

FieldTypeDescription
job_id*stringEchoes the requested job ID.
status*enumOne of pending, processing, completed, failed.
document_type*stringThe requested or detected document type.
page_count*integerNumber of pages in the source document.
resultExtractResponse | nullPresent once status is completed — same shape as the sync ExtractResponse.
errorstring | nullFailure reason, present only when status is failed.
Reference

Errors

Errors are returned as JSON with a consistent shape:

FieldTypeDescription
error*stringA short machine-readable error code.
detailstring | nullHuman-readable detail, when available.
StatusMeaning
400Bad request — malformed body, e.g. neither file nor url provided.
401Missing or invalid bearer token.
402Quota exceeded — either your plan's monthly page quota, or (for free-trial accounts) the shared daily free pool of 7,000,000 tokens/day, which resets at UTC midnight.
415Unsupported file type — only PDF, PNG, and JPEG are accepted.
422Validation error — a field failed schema validation (bad enum value, wrong type, etc.).
429Rate limit exceeded — back off and retry.
502The upstream model provider errored — safe to retry.

A 402 from the shared daily pool carries extra fields alongside error — limit_tokens, used_tokens, and resets_at (ISO timestamp) — so you can show users when to retry instead of just a generic failure.