Skyword Tech — Machine translation and | skywordtech.com

Skyword Tech — Machine translation and | skywordtech.com

How to Test a Machine Translation in 5 Steps

Take 10 segments from your real documents, write your own reference, count the system's errors, then decide: accept, post-edit or reject.

A published BLEU score lives on someone else's test set, and your text is not that set. Ten segments from your actual workload, checked by hand, predict your real experience better than any number on a landing page.

Steps 1 and 2 build the ground truth: pull 10 consecutive segments from real work in about 5 minutes — the typical failure here is cherry-picked clean input — then translate or transcribe them yourself in 10-15 minutes, watching for reference drift, where the reference silently improves on the source.

Steps 3 and 4 produce the count: run the system on the identical input in 2-5 minutes (a different text version than the reference is the classic failure), then mark errors segment by segment for about 10 minutes, keeping normalization consistent — digits versus spelled-out numbers counted the same way on both sides.

The signature mark of this site, drawn as a plate

On narrow screens, swipe or scroll the plate sideways.

Step 5 is the decision. Mark each segment adequate or not: accept if nearly all segments carry the meaning with light fixes, post-edit if errors are systematic but quick to repair, reject if most segments need rewriting. If you want one number, compute BLEU on your 10 segments and keep it as a private baseline for future re-tests.

Budget 30-40 minutes for the whole pass, and repeat it after every model update, because a new version is effectively a new system. The same procedure works for speech-to-text — swap BLEU for word error rate and count substitutions, deletions and insertions instead.

For the transcript side of the check see Word Error Rate: What Counts as Good?, and for reading the score you just computed see BLEU and Other Translation Scores.

  • Step 1 — Sample, 5 min: 10 consecutive segments from real work; failure: cherry-picked input
  • Step 2 — Reference, 10-15 min: translate or transcribe by hand; failure: reference drift
  • Step 3 — Run, 2-5 min: identical input through the system; failure: version mismatch with the reference
  • Step 4 — Count, 10 min: mark errors per segment; failure: inconsistent normalization
  • Step 5 — Decide, 2 min: accept, post-edit or reject against your band

Further reading