Skyword Tech — Machine translation and | skywordtech.com
Word error rate counts substitutions, deletions and insertions against a reference transcript; 3-5% is the professional human bar on clean speech, and anything over 15% needs full rework before use.
The formula is WER = (S + D + I) / total words x 100. On a 20-word reference, one substituted word, one dropped word and one added word make 3 edits — a 15% error rate, or about one wrong word in every seven.
The human bar is the reference point for every band. Professional transcribers run about 3-5% WER on clean read speech and 5-6% on conversational telephone speech, so a machine claim far below 3% usually means an easy test set rather than a better machine.
The bands follow from there. Under 10% the output works as dictation drafts and internal notes; between 10% and 15% it serves internal search and rough summaries but needs post-editing before other people see it; over 15% it goes back for full rework.
On narrow screens, swipe or scroll the plate sideways.
The band also decides whether the speed of speech input survives. Dictation runs at 120-150 words per minute against about 40 on a keyboard — three to four times faster before correction — and a low WER keeps correction time small enough to keep that gain.
Normalize before you count: fix the rules for case, punctuation and digits versus spelled-out numbers, and keep the reference faithful to what was actually said. The same audio scored against two different references yields two different WER figures, so write the denominator down and reuse it across runs.
For translated text the equivalent score is BLEU — see BLEU and Other Translation Scores — and to run the check on your own material follow How to Test a Machine Translation in 5 Steps.
Further reading