Compare vision-language models and locally run OCR engines on the same handwritten pages.
Handwriting accuracy by model.
Choose scripts, benchmark sources, model availability, and CER or WER. The default combines every Latin-script benchmark.
24 models
Latin-script benchmarks
6 benchmarks · 103 pages
Bars show the equal-weight mean across the selected benchmarks. Whiskers are 95% stratified bootstrap confidence intervals from page-level errors. Each page is capped at 100% before averaging; lower is better.
The Pareto frontier.
Orange marks the best accuracy–cost trade-offs. Dashed lines show the error rates of models whose costs are unavailable.
14 models
Pareto frontier Other priced models Price unavailable
Prices are shown per 1,000 public pages, using frozen cost snapshots. The accuracy axis gives each selected benchmark equal weight. Gemini Flash prices use matching token rates; pricing details are noted below.
Compare the models. Review your own result.
This benchmark measures OCR systems on a fixed set of handwritten pages. Your own results may differ.
Upload a page, then check and correct the result before exporting.
Character error rate (CER) counts character substitutions, insertions, and deletions. Word error rate (WER) applies the same edit calculation to whitespace-delimited word tokens, showing how often a reader would encounter a wrong, missing, or extra word. Lower is better for both.
Comparing raw model output can produce nonsense because Markdown table pipes, HTML tags, bullet markers, capitalization, and punctuation can obscure the underlying transcription quality. We therefore converted both output and reference into the same canonical content stream, applied Unicode case-folding, removed Unicode punctuation symmetrically, collapsed the resulting whitespace runs, and then calculated CER and WER. We never used spelling correction, synonym matching, or semantic similarity.
A one-character OCR error
ReferenceMeeting moved to Friday.
ModelMeeting moved to Frlday.
1 substitution ÷ 24 reference characters = 4.2% CER 1 changed token ÷ 4 reference words = 25.0% WER
Presentation normalized
Markdown syntax, HTML formatting, capitalization, and Unicode punctuation do not affect CER or WER.
Structure preserved
Tables are parsed into ordered rows and cells. Missing or malformed table structure remains a separate recorded failure.
Real errors retained
Misspellings, wrong words, missing lines, extra content, truncation, wrong scripts, refusals, and hallucinations remain part of the score.
Dataset-specific rules were narrow, explicit, and applied to both sides.
Examples include folding Bengali digit forms consistently, mapping Hebrew geresh punctuation, and ignoring only the Arabic vowel marks that RASAM’s reference systematically omits. We did not strip accents globally. Two otherwise attractive corpora were removed when their visible content or reading order could not be reconciled safely with their ground truth.
What we tested.
The report separates publisher-provided public datasets from the Inksight Benchmark. The combined selector provides an overview, while every source remains available on its own.
133 public pages
Six datasets × 18 pages plus 25 IAM handwriting-body crops. Each corpus receives equal weight in the public macro CER and WER.
72 Inksight Benchmark pages
Six language and script groups, including 22 English pages, scored against frozen Inksight Benchmark references.
24 reported configurations
All systems received the same frozen prepared images. API models used the byte-identical transcription prompt; local OCR engines used their native pipelines. References and scoring adapters were unchanged.
4,920 reported results
Persistent failures, empty output and unsupported languages received CER and WER 1.0. Inference retries were limited to transient technical failures. Unbilled capacity rejections were reconciled before resuming, not scored as OCR failures.
Public datasets
One row per source used in the final 133-page public benchmark.
Grouped by language and script and scored against frozen Inksight Benchmark references.
Group
Language
Script
Pages
Reference
English · Inksight Benchmark
English
Latin
22
Inksight Benchmark
Arabic · Inksight Benchmark
Arabic
Arabic
10
Inksight Benchmark
Bengali · Inksight Benchmark
Bengali
Bengali
10
Inksight Benchmark
Hindi · Inksight Benchmark
Hindi
Devanagari
10
Inksight Benchmark
Spanish · Inksight Benchmark
Spanish
Latin
10
Inksight Benchmark
French · Inksight Benchmark
French
Latin
10
Inksight Benchmark
Prompt and images
API models were asked to transcribe the document verbatim as UTF-8 GitHub-Flavored Markdown, preserving language, spelling, punctuation, reading order, and visible tables. All systems received the frozen prepared image files. API requests preserved their resolution and format. Local engines then applied their native detection, resizing and recognition steps; these are system comparisons, not identical internal preprocessing.
Pricing reflects snapshots, not live quotes. Gemini 3.7 and 3.8 Flash were discounted when this benchmark was created; both are compared at $0.75 per million input tokens and $3.75 per million output tokens, using recorded token usage. Other models retain their recorded cost snapshots.
We compared base and high reasoning on the same 12 pages across six models, separately from the 205-page benchmark. No model met our criteria for switching to high reasoning, so all retained their base settings.
Does more reasoning help?
Model
CER change (pp)Lower is better
Cost multiplier
Gemini 3 Flash preview
+13.65
8.16×
Gemini 3.1 Pro
+15.94
8.93×
Claude Sonnet 5
+0.23
1.03×
Grok 4.5
-2.59
3.15×
GPT-5.6 Sol
+1.75
2.72×
Qwen 3.8 Max
+2.94
6.27×
High minus base; negative means better. Costs use the experiment’s frozen list prices. These are the original development-set CER scores, not the main benchmark’s normalized, capped scores.
More reasoning did not help on average in this small development-set experiment. Error rates rose by 5.32 percentage points, while costs rose sharply for most models—an average increase of 404% across the six models.