Skyword Tech — Machine translation and | skywordtech.com

Measure it against a reference

Skyword Tech — Machine translation and | skywordtech.com

You measure accuracy by counting errors against a reference you trust. For speech-to-text the score is word error rate: substitutions, deletions and insertions divided by total words, times 100. For machine translation the common score is BLEU, a 0-100 n-gram overlap that only means something inside one test set.

skywordtech.com1 wrong word in 20 at 5% WERBLEU compares only inside one test set

Quality bands in numbers

Word error rate bands and what each one is fit for
Word error rateWhat the output looks likeGood enough forNot good enough for
3-5%One wrong word in 20-33; the professional human level on clean read speechLegal review with a manual check, published transcriptsAlmost nothing fails at this level
5-6%Human level on conversational telephone speechCall analytics, searchable meeting archivesVerbatim legal text without review
Under 10%Occasional wrong word, meaning intactDictation drafts, internal notesBroadcast captions without proofreading
10-15%About one error in every 7-10 wordsInternal search, rough summariesCustomer-facing captions or quotes
Over 15%More than one error in every 7 wordsNothing without full post-editingDirect publication of any kind
Word error rate bands and what each one is fit for

Errors, drawn to scale

Each figure counts the same thing: edits against a reference you wrote yourself.

1966The ALPAC report judges machine translation slower, costlier and less accurate than humantranslation; US government funding for MT falls for roughly a decade.2002BLEU is introduced: a 0-100 geometric mean of n-gram precisions up to 4-grams with a brevitypenalty, still the standard translation score.2017The Transformer paper sets the architecture behind modern systems: 6 encoder and 6 decoderlayers, 512 dimensions, 8 attention heads, about 65 million parameters in the base model and 213million in the big variant.2022Whisper is trained on 680,000 hours of multilingual web audio, a scale no earlier fullysupervised speech corpus matched.2022Meta's No Language Left Behind translates 200 languages — about 3% of the roughly 7,000 livinglanguages.
Sixty years of measuring machines
Word error rateEdit typesReference transcriptBLEU
The key terms of this guide, drawn to one scale

What word error rate actually counts

WER counts every word that was changed, dropped or added, then divides by the words in the reference.

The formula is WER = (S + D + I) / total words x 100. Take a 20-word reference: if the transcript swaps one word, drops another and adds one that was never said, that is 3 edits over 20 words — a 15% error rate, about one wrong word in seven.

All three edit types count, and only those three. A system that invents extra words is penalized exactly like one that skips them, which is why insertions sit in the formula at all.

Before counting, normalize both texts the same way: decide whether case, punctuation and digits versus spelled-out numbers matter. Two people scoring the same audio under different rules will produce two different WER figures.

  • Substitution: the audio says one word, the transcript prints another
  • Deletion: a spoken word is missing from the transcript
  • Insertion: the transcript adds a word nobody said

The method in three parts

Which quality band your task needs

The same 8% transcript is excellent for meeting notes and unacceptable for broadcast captions.

The honest bar is human performance: professional transcribers run about 3-5% WER on clean read speech and 5-6% on conversational telephone speech. A machine claim far below 3% usually signals an easy test set, not a better machine.

Dictation drafts and internal notes stay workable under about 10% — you fix a word here and there. Captions and quotes meant for other people need proofreading once you pass 10%, because one wrong word in ten visibly changes meaning.

Legal review sits at the top of the scale: aim for the human 3-5% band and still verify by hand, because one wrong name or number costs more than the entire check.

What BLEU measures — and what it does not

BLEU scores how closely a machine translation's word sequences match a reference translation, on a 0-100 scale.

Introduced in 2002, BLEU is a geometric mean of n-gram precisions up to 4-grams, with a brevity penalty that punishes translations shorter than the reference. It asks one question: how many of the machine's word sequences also appear in a human translation?

The scale is compressed at the top: on close pairs such as English to French, the best systems sit near 40-50. A score in the 30s can still be usable output, and a 60 can still hide a fatal error in a single sentence.

The hard rule: BLEU compares systems only inside one test set. A 35 on a hard set can beat a 45 on an easy one, so two vendors' published BLEU numbers say almost nothing about each other.

A 30-minute accuracy check on your own sample

Ten segments from your own material tell you more than any published benchmark.

Published benchmarks come from test sets that do not look like your meetings, your accents or your texts. Ten segments from your real workload, checked by hand, predict your actual experience far better.

The decision rule at the end: under 10% WER, accept for drafts; 10-15%, plan to post-edit; over 15%, reject for anything that leaves your desk. The whole check fits in about 30-40 minutes.

  • Step 1 — Sample, 5 min: pull 10 consecutive segments from real work, not a demo file. Typical failure: cherry-picked clean audio.
  • Step 2 — Reference, 10-15 min: transcribe or translate those segments by hand. Typical failure: reference drift, where the reference silently fixes the source.
  • Step 3 — Run, 2-5 min: feed the identical input to the system. Typical failure: a different audio or text version than the reference.
  • Step 4 — Count, 10 min: mark substitutions, deletions and insertions per segment. Typical failure: normalization mismatch — digits versus spelled-out numbers counted inconsistently.
  • Step 5 — Decide, 2 min: compute WER = (S+D+I)/words x 100. Under 10% accept for drafts, 10-15% post-edit, over 15% reject for public use.

Objections that come up every time

Is 95% accuracy the same thing as 5% WER?
Yes — accuracy is just 100 minus WER. The catch is what was counted: check whether insertions made it into the formula, because a score that ignores added words flatters the system.
Can I compare two vendors by their published BLEU scores?
No. BLEU ranks systems only inside a single test set, and two vendors almost never publish on the same one. On close pairs like English-French the best systems sit near 40-50, which shows how compressed the top of the scale is.
Do punctuation and capitalization count as errors?
Only if you decide they do before you start counting. Normalize the reference and the system output under the same rules — case, punctuation, digits versus words — or two people will score the same transcript differently.
Is a 10-segment sample really enough?
It is enough for a first accept, post-edit or reject decision in about 30-40 minutes. Ten segments will not pin down the second decimal of a WER, but they will show which side of the 10% and 15% lines your system sits on.

Where the numbers come from

The timeline above names the primary publication behind each number used on this page.