Choose Inksight
When you want the finished result
- Upload photos, scans, or PDFs
- Review editable text beside the original
- Organize pages and export the document
Handwriting accuracy benchmark · August 2026
We evaluated 166 handwriting pages across two comparisons. Inksight achieved the highest median character accuracy in the 48-page public benchmark and was closest to the text people saved in the separate 118-page production comparison.
Public benchmark
Choose a language to compare the same six systems. Each bar shows median page-level character accuracy; higher is better.
Overall results
Median page accuracy across all 48 public pages
Showing Overall results. Inksight has the highest median character accuracy at 92.8 percent.
Public references. 48 pages with publisher-provided transcripts.
Same input. Identical image bytes and transcription prompt for every system.
Consistent lead. Inksight ranked first overall and in every language shown.
Ranked results
Median page-level character accuracy. Higher is better.Swipe horizontally to compare every language.
| Rank | System | Overall | English | French | Bengali | Arabic |
|---|---|---|---|---|---|---|
| 1 | Inksight | 92.8% | 97.8% | 94.9% | 92.7% | 82.4% |
| 2 | Gemini 3.6 Flash | 91.8% | 96.5% | 94.0% | 91.7% | 79.1% |
| 3 | Claude Opus 5 | 90.5% | 97.1% | 94.3% | 91.0% | 66.0% |
| 4 | Claude Sonnet 5 | 85.8% | 95.1% | 91.7% | 83.3% | 60.4% |
| 5 | GPT-5.6 Terra | 56.4% | 87.3% | 79.2% | 53.0% | 39.7% |
| 6 | GPT-5.6 Luna | 49.3% | 85.6% | 74.7% | 47.9% | 30.3% |
Across two August 2026 evaluations totaling 166 handwriting pages, Inksight ranked first in the 48-page public benchmark and was closest to saved text in the separate 118-page production comparison. Scores are median page-level character accuracy; higher is better. Character accuracy equals 100% minus content character error rate.
Overall, 48 pages. Inksight: 92.78% median character accuracy; Gemini 3.6 Flash: 91.85% median character accuracy; Claude Opus 5: 90.50% median character accuracy; Claude Sonnet 5: 85.83% median character accuracy; GPT-5.6 Terra: 56.41% median character accuracy; GPT-5.6 Luna: 49.27% median character accuracy.
English, 12 pages. Inksight: 97.80% median character accuracy; Gemini 3.6 Flash: 96.46% median character accuracy; Claude Opus 5: 97.06% median character accuracy; Claude Sonnet 5: 95.08% median character accuracy; GPT-5.6 Terra: 87.27% median character accuracy; GPT-5.6 Luna: 85.58% median character accuracy.
French, 12 pages. Inksight: 94.87% median character accuracy; Gemini 3.6 Flash: 93.99% median character accuracy; Claude Opus 5: 94.28% median character accuracy; Claude Sonnet 5: 91.71% median character accuracy; GPT-5.6 Terra: 79.21% median character accuracy; GPT-5.6 Luna: 74.70% median character accuracy.
Bengali, 12 pages. Inksight: 92.66% median character accuracy; Gemini 3.6 Flash: 91.66% median character accuracy; Claude Opus 5: 90.98% median character accuracy; Claude Sonnet 5: 83.28% median character accuracy; GPT-5.6 Terra: 53.01% median character accuracy; GPT-5.6 Luna: 47.85% median character accuracy.
Arabic, 12 pages. Inksight: 82.36% median character accuracy; Gemini 3.6 Flash: 79.08% median character accuracy; Claude Opus 5: 66.03% median character accuracy; Claude Sonnet 5: 60.38% median character accuracy; GPT-5.6 Terra: 39.69% median character accuracy; GPT-5.6 Luna: 30.26% median character accuracy.
Overall median character-accuracy 95% bootstrap intervals: Inksight: 91.10–94.87%; Gemini 3.6 Flash: 89.51–93.21%; Claude Opus 5: 85.41–93.49%; Claude Sonnet 5: 75.84–89.26%; GPT-5.6 Terra: 50.03–69.61%; GPT-5.6 Luna: 44.29–57.77%.
Typical-export set, 36 pages randomly selected from completed exports: Inksight: 94.73% median character accuracy; Claude Opus 5: 79.44% median character accuracy; Claude Sonnet 5: 69.57% median character accuracy; GPT-5.6 Terra: 53.70% median character accuracy; GPT-5.6 Luna: 44.96% median character accuracy.
Known-hard-case set, 82 pages selected because the first transcription was meaningfully corrected: Inksight: 92.27% median character accuracy; Claude Opus 5: 86.25% median character accuracy; Claude Sonnet 5: 78.93% median character accuracy; GPT-5.6 Terra: 63.88% median character accuracy; GPT-5.6 Luna: 57.31% median character accuracy.
Editing-effort audit: 34 of 36 exported pages were kept as-is or changed by less than 1% of characters; two differed by more than 1%.
Scope: this benchmark supports the claim that Inksight outperformed the six systems tested, not that it is universally more accurate than every AI model or on every kind of handwriting.
LLM or specialist OCR?
Claude, GPT, and Gemini can all read handwriting. But a raw model response is only one part of the job. You still need to handle images, prompts, retries, corrections, documents, and exports.
Inksight combines the transcription model with the workflow around it, so you can move from a photo or PDF to editable, exportable text without building an OCR pipeline.
Choose Inksight
Upload your own page and inspect the transcription before exporting.
Editing effort in practice
We also compared five systems with the text saved from 118 production pages across seven languages. Inksight reached 94.7% median character accuracy on 36 randomly selected exports and 92.3% on 82 pages known to need correction.
These saved texts usually began as Inksight output, so this comparison naturally favors Inksight. We use it to show editing effort and difficult cases—not as independent proof of accuracy. The public benchmark above is the fairer model comparison.
Pages people exported
36 pages · Random sample from completed exports
Inksight reached 94.7% median character accuracy on the randomly selected exported pages.
Typical exports: Inksight 94.7%.
Editing-effort audit
In a random sample of completed exports, only two saved transcriptions differed from Inksight’s first result by more than 1% of characters. This suggests little editing before export. It does not prove that every unchanged character was correct, so we report it separately from the public accuracy benchmark.
94.4%
Kept or changed by less than 1%
2 of 36 changed
by more than 1%
Read what our users like and how they use Inksight.
“The best scanning app I have had so far. I really recommend this app—it’s fast and accurate.”
“It is the best scanner and handwritten copier-to-Word app I have ever come across. I recommend it to anyone who needs such an app.”
“A good app—it even deciphered my sloppy handwriting with a minimum of errors.”
“This app is great! So much better than the competition—I had no issue bulk uploading and recognizing my messy handwriting. Recommend.”
“It turns out there are language options. It’s easy to use.”
“Easy to use, and it even recognizes illegible handwriting or crossed-out words on to-do lists. Automatic naming is very helpful, and images of your notes remain visible for reference.”
How accuracy is measured
“Accuracy” does not always mean exact transcription. Some benchmarks measure whether two passages have similar meaning. That can be useful for summarization, but a changed name, date, or number can remain semantically similar while still being wrong.
We use character edits because they more closely represent how much text you would need to correct. We also froze the pages before inference and gave every system the same prepared image and transcription instructions.
Answer key
The publishers supplied the reference text. No Inksight or model-generated text was used as the public benchmark answer key.
Coverage
English, French, Bengali, and Arabic each contribute 12 pages, preventing one language from dominating the sample.
Conditions
Every system received identical prepared image bytes and the same transcription instructions.
Uncertainty
We resampled the 48 pages 20,000 times to show how much the overall result moves with the sample.
We calculate character accuracy as 100% minus content character error rate (CER). CER counts the insertions, deletions, and substitutions needed to match the reference after removing presentation-only Markdown and layout spacing. A score of 92.8% means that about 93 of every 100 characters matched on the median page. It describes the middle page in this sample—not every sentence or every page.
Public source material
Common questions
The useful answer depends on whether you want a model to integrate or a finished tool to use. These answers stay within what the benchmark actually measured.
We evaluated 166 handwriting pages across two August 2026 comparisons. Inksight ranked first among six systems on the 48-page public benchmark, reaching 92.8% median character accuracy overall and ranking first in the English, French, Bengali, and Arabic samples. In a separate 118-page production comparison, Inksight outputs were also closest to the text people saved, although that comparison naturally favors Inksight. These results apply to the systems and pages tested, not every OCR product or type of handwriting.
Gemini 3.6 Flash was the strongest general-purpose model in our test at 91.9% median character accuracy. Claude Opus 5 followed at 90.5%. Inksight, evaluated as a complete handwriting OCR service, ranked above both at 92.8%.
Yes, multimodal GPT models can transcribe handwriting, but accuracy varies considerably by script and page. GPT-5.6 Terra reached 56.4% median character accuracy overall in this multilingual sample, while GPT-5.6 Luna reached 49.3% under the tested configurations.
Yes. Claude Opus 5 reached 90.5% median character accuracy in our public benchmark and was substantially stronger than Claude Sonnet 5 at 85.8%. Both results varied by language, with larger errors on the Arabic pages.
A direct LLM is useful when you are building your own API workflow and want full control over prompts and outputs. A specialist service such as Inksight is the simpler choice when you want to upload pages, review editable text beside the original, organize documents, and export the result without building that workflow yourself.
We report character accuracy as 100% minus content character error rate. The error rate counts the insertions, deletions, and substitutions needed to match the publisher-provided transcript after presentation-only formatting is removed. This is stricter than semantic similarity, which can score text highly even when individual names, numbers, or characters differ.
The test that matters to you
Upload a photo, scan, or PDF and compare the editable transcription with your original.