PRODUCT

Multilingual OCR vision model for Document Intelligence. Powered by ARVA.

April 12, 2026By SAGEA
Back to Index

Document parsing for complex South Asian scripts has long been a major roadblock for digitalization efforts. Today we are introducing ARVA OCR, a high-fidelity, multimodal document parsing system establishing new global benchmarks for printed and handwritten text, built entirely from scratch by SAGEA.

Highlights.

  • Unprecedented Devanagari Accuracy: Outperforms Tesseract v5 and Chandra OCR natively, achieving a staggering 1.2% Character Error Rate (CER) on printed Nepali text.
  • Architecture: A highly optimized 3B-parameter Vision-Language Model (VLM) unified pipeline that natively processes both images and text simultaneously.
  • Training Engine: Powered by SAGEA's proprietary GLUE (Generative Ligature & Union Engine) spatial optimization framework, capable of blending complex conjunct consonants mathematically.
  • Visual Encoding: Dynamic high-DPI resolution patching (supporting up to 4K contextual rendering) to completely mitigate down-sampling losses on dense, small text.
  • Output Topologies: Capable of direct synthesis into structured formats: Markdown for text layouts, LaTeX for complex formulas, and SVG code for rendering charts, diagrams, and physical signatures.

Model Capabilities.

Generalist OCR models frequently collapse when tasked with the conjunct consonants, half-characters, and continuous top-bar lines native to Devanagari and other Indic scripts. Western models rely heavily on disjointed processing pipelines: deploying one model for heuristic bounding-box detection, another for character classification, and brittle post-processing scripts to guess the logical reading order. This results in gibberish when faced with complex, multi-column newspaper layouts or skewed, hand-filled forms.

ARVA operates on a purely monolithic philosophy. By evaluating both text and complex visual components as first-class token sequences in a continuous generative framework, ARVA natively outputs highly structured semantic representations. It fundamentally understands that a caption belongs to an image, that a multi-column newspaper must be parsed sequentially across margins, and that a signature implies a contractual agreement rather than random noise.

The true powerhouse behind ARVA's performance is its integration with SAGEA’s proprietary GLUE framework. GLUE employs intense spatial optimizations—including Poisson boundary blending—to enforce gradient continuity across character junctions during training. This mathematically eliminates the traditional junction seam artifacts that confuse typical vision encoders, allowing ARVA to "read" Indic scripts as a continuous morphological flow rather than a series of disconnected boxes.

Most notably, ARVA demonstrates an unprecedented 5.8% Character Error Rate on heavily degraded, organically hand-filled Nepali forms. By expanding the native VLM tokenizer vocabulary to faithfully represent the rich morphological structure of the Devanagari script, ARVA outright sidesteps the sub-word segmentation and spatial failures frequently observed when imposing Western OCR parsers onto Indic topography.

Nepali Document Benchmark
Complex Tables Structural Accuracy (%) ↑
ARVA
92.4
Chandra OCR
84.5
Capability BenchmarkARVA OCRChandra OCRTesseract v5
Nepali Printed Text (CER)1.2%4.6%18.5%
Nepali Handwritten Text (CER)5.8%12.1%>60.0%
Mixed Script Layout (Nepali/EN)95.1%82.8%32.1%

*Note: Lower Character Error Rate (CER) is superior. Benchmarks conducted at int8 precision on a highly curated dataset of historical documents, physical municipal records, and bank ledgers.

Use Cases.

  • Historical Archive Digitization: Rapidly process and digitize decades of degraded, physical municipal records, land deeds, and classical literature natively into highly structured, searchable Markdown databases without losing layout context.
  • Automated KYC Processing: Deploy ARVA natively into banking and fintech workflows to instantly parse, verify, and structure data from heavily artifacted scans of physical South Asian identity cards (e.g. Citizenship cards, Passports) with near-perfect reliability.
  • Enterprise Invoice Parsing: Automatically ingest multi-page, mixed-script enterprise invoices and tax forms. ARVA can reconstruct complex financial tables flawlessly, turning unstructured PDFs into clean JSON data ready for backend ERP systems.

Get started.

ARVA OCR is available today in SAGEA Studio and APIs, and powers remote coding agents and Work mode on the Pro, Team, and Enterprise plans.

It is available for prototyping and production deployment, hosted on SAGEA-accelerated endpoints on platform.sagea.space.

Build the future of agentic systems with us.

We're hiring across research, engineering, and product to push agentic systems further. See our open roles.