Is BLEU the best way to judge translation quality?
If you've ever compared translation systems, you've probably seen BLEU scores thrown around as the gold standard. But here's the thing: BLEU was introduced by IBM researchers back in 2002, and it's increasingly seen as unreliable because it correlates poorly with human judgment (IBM Research). Yet many people still lean on it because it's cheap and familiar. I think that's a mistake. We have better tools now, and we should use them.
Why does BLEU still dominate discussions?
BLEU is easy to compute—you just compare n-grams between machine output and a human reference. That simplicity made it the default for two decades. But as the IBM Research blog points out, newer metrics like COMET and BLEURT are becoming standard because they align better with what humans actually think is good. The problem is that BLEU is so entrenched that many researchers and vendors still quote it as if it's meaningful. I've seen products claim 'high BLEU' as if that guarantees quality—but it doesn't.
What's wrong with BLEU in practice?
BLEU measures n-gram overlap, not meaning. A translation can be perfectly understandable but use different words, and BLEU will penalize it unfairly. Conversely, a literal but awkward translation might score higher because it shares more words with the reference. That's a fundamental flaw. For example, in patent translation, where terminology is crucial, BLEU might favor a translation that matches the reference terms but reads poorly. Meanwhile, a human would prefer a more natural phrasing that BLEU marks down.
What are the alternatives to BLEU?
There are several. METEOR, from Carnegie Mellon, uses both precision and recall, and handles synonyms and paraphrases (METEOR). chrF, proposed in 2015, works on character n-grams and is great for morphologically rich languages (chrF). TER measures how many edits a human would need to make to the machine output to match a reference—it's a direct proxy for post-editing effort (Translation Edit Rate). But the most exciting are neural metrics like COMET and BLEURT. COMET, developed by Unbabel, won the WMT shared tasks in 2019 and 2020 (COMET). BLEURT, from Google, is BERT-based and correlates better with human judgment than BLEU on several tasks (BLEURT). These are not perfect, but they're a huge step up.
Does BLEU still have any use?
Sure, for quick sanity checks during development. The Transformer paper reported a BLEU score of 28.4 on English-to-German, improving on the previous best by more than 2 points (Attention Is All You Need). That kind of relative comparison can be useful when you're iterating on a model. But for evaluating whether a translation is actually good enough to publish, BLEU is not enough. I'd argue that if you're using BLEU to compare competing systems for a client, you're doing them a disservice.
What about human evaluation?
Human evaluation is the gold standard, but it's expensive and slow. That's why researchers have been trying to create automatic metrics that approximate human judgment. For instance, the WMT metrics shared task assesses MT quality with and without references, while quality estimation does it without any reference (WMT22). So the field is moving toward metrics that don't need a reference at all. But until those are perfect, we should use the best available proxies—and that means moving beyond BLEU.
So what should you actually use?
My recommendation is simple: stop relying on BLEU alone. If you're a researcher, report COMET or BLEURT alongside BLEU. If you're a buyer of translation services, ask for human evaluation or at least for metrics that correlate with human judgment. The language industry is huge—CSA Research estimated the market at US$49.68 billion in 2023 (CSA Research)—and you deserve better than a metric from 2002 that was never designed to measure quality. Let's demand more.
The takeaway
BLEU is a relic that oversimplifies translation quality and misleads us. Newer metrics like COMET and BLEURT are not perfect, but they're clearly better. If you care about translation quality—whether you're a translator, a project manager, or a researcher—you should push for evaluation methods that reflect human preferences. The future of translation evaluation is here, and it's time to let go of BLEU.
Sources
- IBM Research - https://research.ibm.com/blog/bleu-nlp-benchmark-anniversary
- METEOR - http://www.cs.cmu.edu/~alavie/METEOR/
- chrF - https://aclanthology.org/W15-3049/
- COMET - https://aclanthology.org/2020.emnlp-main.213/
- Translation Edit Rate - https://aclanthology.org/2006.amta-papers.25/
- BLEURT - https://research.google/pubs/bleurt-learning-robust-metrics-for-text-generation/
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!