What's the best way to fix machine translation that keeps producing garbage? You've probably tried better MT engines, more post-editing, or different metrics. But here's the blunt answer: stop obsessing over BLEU scores and start applying translation theory.
The Myth of the Perfect Metric
You've been told that BLEU is the gold standard. It was introduced by IBM researchers in 2002, and it compares n-gram matches between machine output and human references. It's simple. But it's also increasingly seen as unreliable because it correlates poorly with human judgment (IBM Research). So you keep chasing a higher BLEU, thinking it means better quality. Meanwhile, your translators are rolling their eyes because the MT still misses the cultural nuance that a human would catch. Translation theory, the old debates about equivalence and skopos, offers a better path.
What Translation Theory Actually Gives You
Translation theory isn't just academic navel-gazing. It gives you a framework to decide what "good" means for your project. For example, the concept of skopos—the purpose of the translation—tells you that a legal contract and a marketing slogan require different strategies. When you use MT, you need to know the purpose, or you'll end up with a literal, unusable draft. ISO 17100:2015, the standard for translation services, expects you to consider the purpose too. And ISO 18587:2017 for post-editing of MT output stresses the need for human competences that go beyond fixing grammar. So, when you set up your post-editing workflow, you're not just checking for errors—you're making theoretical decisions about whether the MT output fits the purpose.
Better Metrics? Only If You Use Them Right
Now, you might argue: "But metrics like COMET and BLEURT are better than BLEU." True, they correlate better with human judgment (IBM Research). But they still measure surface-level quality. They can't tell you if the translation is culturally appropriate or if it reads naturally in the target language. That's where theory comes in. You need to train your post-editors to think theoretically: to ask, "What is the function of this text?" "What does the target audience expect?" "What are the norms of this genre?" For instance, if you're translating Japanese to Korean with zero-shot MT (which Google's multilingual NMT made possible in 2016), you might get a perfectly grammatical output that sounds like a machine. Theory tells you to check for idiomaticity, not just fluency.
How to Apply This Now
So, here's your action plan. Stop reporting BLEU as your quality KPI. Instead, define quality based on the translation's purpose, using the skopos theory. Set up a post-editing workflow that includes a brief from a human expert who knows the target audience. Use metrics like TER or HTER to measure post-editing effort, since HTER uses a human reference closest to the MT output (Translation Edit Rate). But don't rely on them alone. Have your translators annotate why they made changes—was it a factual error, a stylistic issue, or a cultural adaptation? That feedback loop will improve your MT training data and your prompts.
- Adopt a "purpose-first" mindset: define the function before you hit translate.
- Use ISO 18587:2017 as a guide for post-editing competences.
- Track HTER, not BLEU, to measure actual rework.
You might say: "But this is too much work. We just need a quick fix." I get it. But the quick fix—chasing BLEU—has failed you. The 1966 ALPAC report cut MT funding because the field overpromised and underdelivered. Don't repeat that mistake. The language services market is shrinking (CSA Research estimated $49.68 billion in 2023, down 4.5% from the year before). To stay competitive, you need to deliver real quality, not just a number.
Bottom line
Your single best move: ditch BLEU as your primary metric and adopt a purpose-driven approach to MT post-editing. Define the skopos, use HTER to measure effort, and train your team to think like translators, not just editors. That's how you turn MT from a liability into an asset.
Sources
- IBM Research - https://research.ibm.com/blog/bleu-nlp-benchmark-anniversary
- ISO 18587:2017 - https://www.iso.org/standard/62970.html
- Translation Edit Rate - https://aclanthology.org/2006.amta-papers.25/
- ALPAC - https://en.wikipedia.org/wiki/ALPAC
- CSA Research - https://csa-research.com/l/media/Language-Services-and-Technology-Industry-Faces-Revenue-Decline-but-Remains-Poised-for-Transformation
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!