Skip to main content
Localization

Stop Worshiping BLEU: Localization Needs Human Judgment

BLEU scores mislead localization teams. Human evaluation and post-editing standards like ISO 18587 should drive quality decisions, not automated metrics.

BLEU Is a Crutch, Not a Metric

Everyone in localization quotes BLEU scores like they're sacred. But BLEU is a broken compass. It was introduced by IBM researchers in 2002 to measure machine translation quality, and it's still used everywhere, despite the fact that it correlates poorly with human judgment (IBM Research). BLEU compares n-grams between machine output and a reference translation, and it rewards exact matches. That sounds fine, but it fails to capture meaning, nuance, and fluency. A translation can be perfect and still score low if the human reference used different wording. BLEU was never designed for the messy reality of localization, where context, tone, and cultural fit matter more than word overlap. Relying on BLEU to make decisions about your localization pipeline is like judging a chef by how many times they use the word 'salt' in their recipes.

Localization Is Not a Math Problem

Localization is about adapting a product or content to a specific market. It involves idioms, cultural references, and domain-specific terminology. Machine translation (MT) has come a long way, especially with neural networks. The Transformer architecture, introduced in 2017, became the foundation of modern neural machine translation, and systems like Google Translate and DeepL are widely used (Machine translation, Wikipedia). DeepL has reported that in blind tests with professional translators, its translations were chosen roughly three times more often than those of Google, Microsoft, or Facebook (DeepL Translator, Wikipedia). That's impressive, but it doesn't mean MT can replace a human. MT output often requires human post-editing, especially for idiomatic expressions, cultural nuance, and domain-specific terminology (Machine translation, Wikipedia). That's where the real work lies.

The Counterargument: 'But Automation Saves Time and Money'

I hear that from every vendor. They say, 'We use MT and post-editing, and we save 50% on costs.' Maybe. But what are you sacrificing? Quality. And when you cut corners on quality, you lose customer trust. The European Commission's Directorate-General for Translation, one of the largest translation services in the world, translates into and out of all 24 official EU languages and produced about 2.6 million translated pages in 2022 (European Commission translation department). They don't rely on raw MT for those pages. They have strict standards and human oversight. The ISO 17100:2015 standard specifies requirements for core processes, resources, and other aspects of delivering a quality translation service (ISO 17100:2015). And ISO 18587:2017 specifies requirements for full human post-editing of MT output (ISO 18587:2017). These standards exist for a reason. If you're skipping human review, you're not doing professional localization.

What Should You Use Instead of BLEU?

Better metrics exist. COMET, developed by Unbabel, is a neural framework that achieved top performance at WMT 2019 and WMT 2020 metrics shared tasks (COMET, ACL Anthology). BLEURT, from Google Research, correlates better with human judgments than BLEU on several natural language generation tasks (BLEURT, Google Research). These are not perfect, but they're a step up. Even more important is human evaluation. Use professional translators to assess quality, or at least use HTER (Human-targeted Translation Edit Rate), which measures how much a human needs to edit the MT output (Translation Edit Rate, ACL Anthology). That gives you a direct measure of post-editing effort, which is what really matters for your budget.

What I'd Actually Do

Stop using BLEU as your primary quality gate. If you're a localization manager, set up a two-stage process: first, run MT output through a strong neural metric like COMET or BLEURT to flag likely problem segments. Then, have a human post-editor review those flagged segments, following ISO 18587:2017 guidelines. For high-stakes content, skip the automated gate entirely and have a professional translator do full post-editing. Invest in building a robust translation memory (TM) database, because reusing high-quality human translations is always better than starting from MT (Machine translation, Wikipedia). And when you evaluate MT vendors, don't ask for BLEU scores. Ask for human preference tests, like DeepL's blind tests with professional translators (DeepL Translator, Wikipedia). That's the kind of evidence that matters. In the end, localization is about people, not algorithms. Act like it.

Sources

  • IBM Research - https://research.ibm.com/blog/bleu-nlp-benchmark-anniversary
  • DeepL Translator - https://en.wikipedia.org/wiki/DeepL_Translator
  • European Commission translation department - https://commission.europa.eu/about-european-commission/departments-and-executive-agencies/translation_en
  • ISO 18587:2017 - https://www.iso.org/standard/62970.html
  • COMET - https://aclanthology.org/2020.emnlp-main.213/
  • BLEURT - https://research.google/pubs/bleurt-learning-robust-metrics-for-text-generation/

Share this article:

Comments (0)

No comments yet. Be the first to comment!