Skyword Tech — Machine translation and | skywordtech.com
BLEU is a 0-100 measure of n-gram overlap with a reference translation; it ranks systems inside one test set, and the best systems on close pairs such as English-French sit near 40-50.
BLEU, introduced in 2002, is a geometric mean of n-gram precisions up to 4-grams between a machine translation and a human reference, with a brevity penalty that punishes output shorter than the reference. It counts matching word sequences, not understanding.
The scale is compressed at the top. On close pairs such as English to French the strongest systems sit near 40-50, a score in the 30s can still be usable output, and a 60 can still contain a fatal error in one critical sentence.
The comparability rule is absolute: BLEU ranks systems only inside a single test set. A 35 on a hard set can beat a 45 on an easy one, so two vendors quoting BLEU from different sets are not comparable in either direction.
On narrow screens, swipe or scroll the plate sideways.
chrF applies the same idea to character n-grams instead of words, which helps where word boundaries blur or morphology is rich, and the same one-test-set rule applies. Human adequacy and fluency scales take the opposite approach: raters judge whether the meaning survived, which is slower but catches what n-gram overlap misses.
Modern scores sit on modern architecture: the 2017 Transformer paper built its base model from 6 encoder and 6 decoder layers at 512 dimensions with 8 attention heads — about 65 million parameters, 213 million in the big variant — and later systems scaled from there. The judging rule, however, has not moved.
For spoken input the parallel score is word error rate — see Word Error Rate: What Counts as Good? — and to produce your own number instead of trusting a published one, follow How to Test a Machine Translation in 5 Steps.
Further reading