Skip to main content
Localization

Stop Obsessing Over BLEU: A Localization Editor's Guide to MT Quality

Forget the BLEU score. If you're localizing, the real metric is post-editing effort. Here's how to set up a workflow that actually measures what matters.

Who This Is For

If you're a localization manager, a freelance translator who's been asked to "just run it through Google Translate," or a product owner who's suddenly responsible for shipping your app in five languages, this guide is for you. I've spent years in the trenches of translation, and I've seen too many teams make the same mistake: they obsess over the wrong numbers. They chase a high BLEU score, pat themselves on the back, and then ship a translation that reads like it was generated by a sleep-deprived robot. I'm here to tell you: stop it. BLEU is a crutch, and it's leading you astray.

Why BLEU Is a Crutch

BLEU, the Bilingual Evaluation Understudy, was introduced by IBM researchers back in 2002 as a quick, automated way to compare machine translation output to human references. It works by counting n-gram overlaps — basically, how many words and phrases match between the machine output and a human translation. And for a long time, it was the standard. But here's the dirty secret: BLEU correlates poorly with human judgment. It's been known for years that a high BLEU score doesn't mean a translation is good; it just means it's similar to the reference. As IBM Research itself has noted, BLEU is increasingly seen as unreliable, and newer metrics like COMET and BLEURT are becoming the standard. So why are we still using it to make decisions about localization? Because it's easy. But easy isn't good enough when you're spending real money on human post-editors.

Step 1: Define What "Good" Means for Your Project

Before you even look at machine translation, sit down and define what quality means for your specific content. Is it a marketing email where tone and wit matter? Or is it a safety manual where precision is non-negotiable? The European Commission's Directorate-General for Translation, one of the largest translation services in the world, translates into and out of 24 official languages and produced about 2.6 million pages in 2022 alone. They don't rely on a single metric; they have human reviewers at every step. You need to think like them. For most content, you'll want a mix of automatic evaluation and human review. But the key is to define your quality bar *before* you start, not after.

Step 2: Pick the Right MT Engine (and Test It Yourself)

Don't just default to the biggest name. DeepL, Google, Microsoft, Amazon — they all have strengths and weaknesses. DeepL, launched in 2017, has claimed that in blind tests with professional translators, its translations were chosen roughly three times more often than Google's, Microsoft's, or Facebook's. And that might be true for European languages. But for low-resource languages, Meta's NLLB-200 model, which supports 200 languages, improved translation quality by an average of 44% compared to previous systems, with gains exceeding 70% for some African and Indian languages. The point is: test on *your* content, in *your* language pair. Don't trust a marketing blog. Run a sample through two or three engines and have a human evaluate the output — not with BLEU, but with their eyes.

Step 3: Use a CAT Tool to Keep Humans in the Loop

Machine translation is not a replacement for human translators; it's a tool to make them more efficient. That's the whole idea behind computer-assisted translation (CAT). Tools like SDL Trados, memoQ, or even free ones like OmegaT let you use translation memory (TM) — a database of previously translated segments — alongside MT. The Google Translator Toolkit, which ran from 2009 to 2019, worked exactly this way: it split documents into segments, pre-translated them from TM, and only fell back on MT when no match was found. That's the right approach. You want your translator to see the best possible starting point, whether that's a perfect TM match or a raw MT suggestion. And when they do use MT, they're post-editing, not translating from scratch. That's where the real cost savings come in.

Step 4: Measure Post-Editing Effort, Not BLEU

Here's the radical idea: stop looking at BLEU entirely for your localization decisions. Instead, measure how much effort it takes a human to fix the MT output. There's a metric for that: Translation Edit Rate (TER), which counts the number of edits — insertions, deletions, substitutions, and shifts — needed to change the MT output into a reference translation. A variant called HTER even uses a human-created reference that's closest to the MT output, making it a better proxy for post-editing effort. But you don't need to get that formal. Just track the time your post-editors spend on each segment. If they're spending 10 minutes on a 100-word paragraph, that's a red flag. If they're spending 30 seconds, you're golden. That's the number that matters for your budget, your timeline, and your sanity.

Step 5: Invest in Post-Editing Standards and Training

Post-editing isn't just "fixing a few words." It's a skill. ISO 18587:2017 sets requirements for full human post-editing of MT output, including the competences a post-editor must have. And ISO 17100:2015 does the same for translation services overall. If you're going to use MT, you need to train your team on how to post-edit effectively. That means knowing when to fix a sentence structure, when to reorder a phrase for naturalness, and when to completely rewrite a segment because the MT output is garbage. And it means knowing the difference between light post-editing (just making it understandable) and full post-editing (making it publishable). Don't assume your translators know how to do this; train them explicitly.

What Can Go Wrong: The Zero-Shot Trap

Here's a warning: don't get seduced by zero-shot translation. In 2016, Google announced that its multilingual NMT system could translate between language pairs it had never been trained on, like Japanese to Korean. It was a breakthrough, and it hinted at a kind of universal "interlingua" in the model. But zero-shot translation is still fragile. It works best for closely related languages, and it can produce catastrophic errors for distant pairs. I've seen a zero-shot translation from Finnish to Swahili that was pure gibberish. So, use zero-shot with caution, and always have a human review the output for high-stakes content. The same goes for any MT output, really. No matter how good the engine is, it will make mistakes, especially with idiomatic expressions, cultural nuance, and domain-specific terminology. That's just the nature of the beast.

Bottom Line

The single best move you can make is to stop caring about BLEU and start measuring post-editing effort. It's the only metric that directly reflects your real costs and your real quality. So, set up a simple workflow: pick your MT engine based on your own tests, integrate it into a CAT tool, track how long your post-editors spend on each segment, and invest in training them to do it well. Do that, and you'll ship translations that are actually good — not just statistically similar.

Sources

  • Machine translation (Wikipedia) - https://en.wikipedia.org/wiki/Machine_translation
  • IBM Research - https://research.ibm.com/blog/bleu-nlp-benchmark-anniversary
  • DeepL Translator (Wikipedia) - https://en.wikipedia.org/wiki/DeepL_Translator
  • No Language Left Behind (Nature) - https://www.nature.com/articles/s41586-024-07335-x
  • Translation Edit Rate (ACL Anthology) - https://aclanthology.org/2006.amta-papers.25/
  • ISO 18587:2017 (ISO) - https://www.iso.org/standard/62970.html

Share this article:

Comments (0)

No comments yet. Be the first to comment!