Handwritten text extraction benchmark

Compare vision-language models and locally run OCR engines on the same handwritten pages.

Handwriting accuracy by model.

Choose scripts, benchmark sources, model availability, and CER or WER. The default combines every Latin-script benchmark.

24 models

Latin-script benchmarks

6 benchmarks · 103 pages

Average normalized CER by modelLower bars are better. Whiskers show 95% stratified bootstrap confidence intervals across the selected benchmark pages.0%10%20%30%40%50%60%70%Claude Opus 5: 8.2% CER, 95% CI 6.3–10.4%8.2Opus 5Gemini 3.7 Flash: 8.4% CER, 95% CI 6.5–10.7%8.4Gemini 3.7 FlashGrok 4.5: 8.6% CER, 95% CI 7.0–10.3%8.6Grok 4.5Gemini 3 Flash: 9.1% CER, 95% CI 6.5–12.0%9.1Gemini 3 FlashGemini 3.8 Flash: 9.1% CER, 95% CI 7.1–11.5%9.1Gemini 3.8 FlashQwen 3.8 Max: 9.2% CER, 95% CI 7.3–11.3%9.2Qwen 3.8 MaxClaude Sonnet 5: 9.3% CER, 95% CI 7.4–11.4%9.3Sonnet 5Gemini 3.1 Pro: 9.3% CER, 95% CI 7.1–11.8%9.3Gemini 3.1 ProGPT-6 Astra: 9.8% CER, 95% CI 7.6–12.3%9.8GPT-6 AstraGemini 3.5 Flash Lite: 10.3% CER, 95% CI 8.2–12.6%10.3Gemini 3.5 LiteKimi K3: 10.8% CER, 95% CI 8.3–13.4%10.8Kimi K3GPT-5.6 Sol: 12.1% CER, 95% CI 10.1–14.4%12.1GPT-5.6 SolGLM 5.3 Flash: 12.1% CER, 95% CI 9.8–14.6%12.1GLM 5.3 FlashQwen 3.8 27B: 13.0% CER, 95% CI 10.7–15.6%13.0Qwen 3.8 27BGPT-5.6 Terra: 15.4% CER, 95% CI 12.1–19.3%15.4GPT-5.6 TerraGPT-5.6 Luna: 16.3% CER, 95% CI 13.5–19.9%16.3GPT-5.6 LunaGemma 4 26B-A4B: 16.4% CER, 95% CI 13.2–20.0%16.4Gemma 4 26B-A4BQwen 3.7 Flash: 18.2% CER, 95% CI 13.8–22.7%18.2Qwen 3.7 FlashLightOnOCR-2: 20.9% CER, 95% CI 15.7–26.3%20.9LightOnOCR-2GLM-OCR: 26.1% CER, 95% CI 21.4–30.5%26.1GLM-OCRPaddleOCR-VL-1.6: 27.4% CER, 95% CI 22.6–32.5%27.4PaddleOCR-VL-1.6DeepSeek V4.1 Flash: 32.4% CER, 95% CI 26.1–39.1%32.4DeepSeek V4.1 FlashEasyOCR 1.7.2: 58.0% CER, 95% CI 55.2–61.0%58.0EasyOCR 1.7.2Tesseract 5.5.0: 65.6% CER, 95% CI 62.3–68.8%65.6Tesseract 5.5LOWER IS BETTER · ORANGE = LOWEST ERROR · 95% CI103 pages · equal weight per benchmark

Bars show the equal-weight mean across the selected benchmarks. Whiskers are 95% stratified bootstrap confidence intervals from page-level errors. Each page is capped at 100% before averaging; lower is better.

The Pareto frontier.

Orange marks the best accuracy–cost trade-offs. Dashed lines show the error rates of models whose costs are unavailable.

14 models
Pareto frontier Other priced models Price unavailable
Cost per 1,000 pages versus normalized CERLower cost and lower error are better. Highlighted models form the frontier for the selected benchmarks.CER0%10%20%30%$0.1$1$10$100Qwen 3.7 FlashGemini 3.5 LiteGemini 3 FlashGemini 3.7 FlashOpus 5LOWER COST · LOWER ERROR IS BETTERCost per 1,000 pages (USD, logarithmic)

Prices are shown per 1,000 public pages, using frozen cost snapshots. The accuracy axis gives each selected benchmark equal weight. Gemini Flash prices use matching token rates; pricing details are noted below.

Compare the models. Review your own result.

This benchmark measures OCR systems on a fixed set of handwritten pages. Your own results may differ.

Upload a page, then check and correct the result before exporting.

Try it

Notes on preprocessing and error calculation

Character error rate (CER) counts character substitutions, insertions, and deletions. Word error rate (WER) applies the same edit calculation to whitespace-delimited word tokens, showing how often a reader would encounter a wrong, missing, or extra word. Lower is better for both.

Comparing raw model output can produce nonsense because Markdown table pipes, HTML tags, bullet markers, capitalization, and punctuation can obscure the underlying transcription quality. We therefore converted both output and reference into the same canonical content stream, applied Unicode case-folding, removed Unicode punctuation symmetrically, collapsed the resulting whitespace runs, and then calculated CER and WER. We never used spelling correction, synonym matching, or semantic similarity.

A one-character OCR error

ReferenceMeeting moved to Friday.
ModelMeeting moved to Frlday.
1 substitution ÷ 24 reference characters = 4.2% CER
1 changed token ÷ 4 reference words = 25.0% WER

Presentation normalized

Markdown syntax, HTML formatting, capitalization, and Unicode punctuation do not affect CER or WER.

Structure preserved

Tables are parsed into ordered rows and cells. Missing or malformed table structure remains a separate recorded failure.

Real errors retained

Misspellings, wrong words, missing lines, extra content, truncation, wrong scripts, refusals, and hallucinations remain part of the score.

Dataset-specific rules were narrow, explicit, and applied to both sides.

Examples include folding Bengali digit forms consistently, mapping Hebrew geresh punctuation, and ignoring only the Arabic vowel marks that RASAM’s reference systematically omits. We did not strip accents globally. Two otherwise attractive corpora were removed when their visible content or reading order could not be reconciled safely with their ground truth.

What we tested.

The report separates publisher-provided public datasets from the Inksight Benchmark. The combined selector provides an overview, while every source remains available on its own.

133 public pages

Six datasets × 18 pages plus 25 IAM handwriting-body crops. Each corpus receives equal weight in the public macro CER and WER.

72 Inksight Benchmark pages

Six language and script groups, including 22 English pages, scored against frozen Inksight Benchmark references.

24 reported configurations

All systems received the same frozen prepared images. API models used the byte-identical transcription prompt; local OCR engines used their native pipelines. References and scoring adapters were unchanged.

4,920 reported results

Persistent failures, empty output and unsupported languages received CER and WER 1.0. Inference retries were limited to transient technical failures. Unbilled capacity rejections were reconciled before resuming, not scored as OCR failures.

Public datasets

One row per source used in the final 133-page public benchmark.

DatasetMaterialLanguage · scriptPagesLicenseProject / paper
BN-HTRdModern handwritten proseBengali · Bengali18CC BY 4.0Source
ScaDS GermanModern copied proseGerman · Latin18CC BY 4.0Source
RASAMHistorical manuscript proseArabic · Arabic18Apache 2.0Source
PinkasHistorical community ledgerHebrew · Hebrew18CC BY 4.0Source
EPARCHOSHistorical codexHistorical Greek · Greek18CC BY 4.0Source
POPP CensusHistorical census tablesFrench · Latin18CC BY 4.0Source
IAM HandwritingModern handwritten copied proseEnglish · Latin25IAM non-commercial research termsSource

Inksight Benchmark

Grouped by language and script and scored against frozen Inksight Benchmark references.

GroupLanguageScriptPagesReference
English · Inksight BenchmarkEnglishLatin22Inksight Benchmark
Arabic · Inksight BenchmarkArabicArabic10Inksight Benchmark
Bengali · Inksight BenchmarkBengaliBengali10Inksight Benchmark
Hindi · Inksight BenchmarkHindiDevanagari10Inksight Benchmark
Spanish · Inksight BenchmarkSpanishLatin10Inksight Benchmark
French · Inksight BenchmarkFrenchLatin10Inksight Benchmark

Prompt and images

API models were asked to transcribe the document verbatim as UTF-8 GitHub-Flavored Markdown, preserving language, spelling, punctuation, reading order, and visible tables. All systems received the frozen prepared image files. API requests preserved their resolution and format. Local engines then applied their native detection, resizing and recognition steps; these are system comparisons, not identical internal preprocessing.

Pricing reflects snapshots, not live quotes. Gemini 3.7 and 3.8 Flash were discounted when this benchmark was created; both are compared at $0.75 per million input tokens and $3.75 per million output tokens, using recorded token usage. Other models retain their recorded cost snapshots.

Does more reasoning help?

We compared base and high reasoning on the same 12 pages across six models, separately from the 205-page benchmark. No model met our criteria for switching to high reasoning, so all retained their base settings.

Does more reasoning help?
ModelCER change (pp)Lower is betterCost multiplier
Gemini 3 Flash preview+13.658.16×
Gemini 3.1 Pro+15.948.93×
Claude Sonnet 5+0.231.03×
Grok 4.5-2.593.15×
GPT-5.6 Sol+1.752.72×
Qwen 3.8 Max+2.946.27×

High minus base; negative means better. Costs use the experiment’s frozen list prices. These are the original development-set CER scores, not the main benchmark’s normalized, capped scores.

More reasoning did not help on average in this small development-set experiment. Error rates rose by 5.32 percentage points, while costs rose sharply for most models—an average increase of 404% across the six models.