Skip to main content
Insightech
3 min read

Vietnamese OCR: why accuracy is not a single number

The question "how accurate is your OCR" sounds reasonable but has no correct answer. Here is why, and what to ask instead.

  • OCR
  • digitisation
  • evaluation

In nearly every conversation about document digitisation, one question comes up: “How accurate is your recognition software?”

It is a fair question. But the honest answer is that there is no single number that transfers between situations. This article explains why, and suggests what to evaluate instead.

The same software, wildly different results

Consider three documents fed into the same system:

A laser-printed decision, scanned at 300 DPI, standard typeface. Results are close to perfect. Errors, if any, are limited to a few unusual characters.

A dispatch photocopied several times, slightly faded, with a stamp overlapping the text. Results drop noticeably. The text under the stamp is usually lost.

An archived file from a decade ago, yellowed paper, with handwritten sections. The printed text still reads acceptably; the handwritten parts are inconsistent.

Same software, three very different quality levels. So which number is “the software’s accuracy”? None of them — it is the software’s accuracy on that type of document.

Why Vietnamese is harder

Vietnamese has characteristics that make recognition considerably harder than English.

Diacritics stack. A single character can carry both a vowel modifier and a tone mark. On a blurred or low-resolution scan these marks are easily lost or merged. Losing a tone mark changes the meaning of the word entirely.

Many word pairs differ only by a diacritic. Misreading a mark does not produce obvious nonsense; it produces a different word that is also valid. This is the dangerous class of error, because a standard spellchecker will not catch it.

Administrative documents contain distinctive notation. Reference numbers in the form “Số: 123/QĐ-UBND” mix letters, digits and slashes in patterns that general-purpose recognition models have not seen much of.

What to evaluate instead

Rather than asking for a number, we suggest this approach.

One, supply a real sample. Select fifty to a hundred pages that genuinely reflect your archive — meaning both the clean documents and the poor ones, in their real proportions. Do not submit only the cleanest material, because the measured result will not reflect operational reality.

Two, measure on that sample. The vendor runs the trial and reports results on your sample, broken down by document quality group.

Three, ask what happens when it is unsure. This is the most important question and the one least often asked: when the system recognises text with low confidence, what does it do? There are two possibilities, and they are very far apart.

The question that matters most

A system that is 95% accurate and clearly flags the other 5% as uncertain is usable. Officers know where to check.

A system that is 97% accurate but presents all 100% identically is considerably more dangerous. Three percent of errors are mixed in among the correct text, nobody knows where they are, and they quietly enter the archive to be cited in some report later.

For administrative records, a system’s ability to recognise its own limits matters as much as raw accuracy. That is why we design for the first behaviour: low-confidence pages are flagged for an officer to re-check, rather than silently accepted.

Preparing documents before scanning

A few simple measures improve results substantially, and doing them before scanning is far cheaper than fixing afterwards:

  • Scan at 300 DPI or above. Below that, Vietnamese diacritics start disappearing.
  • Scan black-and-white for ordinary printed text, but keep colour where a red stamp overlaps the text.
  • Flatten folded documents before scanning; creases cast shadows and destroy the line of text under them.
  • Separate poor-quality material into its own batch with its own handling, rather than mixing it in.

In short

If a vendor gives you an accuracy figure before seeing your documents, that figure rests on nothing. The only way to know whether the software works on your archive is to run it against your archive.


We run trials on real samples from your organisation and report the measured results, including the document groups the system handled poorly.

See it run on your own documents

Every solution sounds good in a description. The only way to know whether this one works for you is to run it against your real documents, templates and workflows. That is exactly the kind of demo we do.

The demo is free and carries no obligation. If it turns out we are not the right fit, we will say so.