Contrarian Hook: MT Is Not Ready for Prime Time
You've heard it a thousand times: neural machine translation is 'good enough' to skip the human. But if you're a localization manager or a freelance translator, you know that's a trap. The fact is, MT output still needs human post-editing—especially for idiomatic expressions, cultural nuance, and domain-specific terminology (Machine translation, Wikipedia). And the standards agree: ISO 18587:2017 exists specifically to define what full human post-editing of MT output should look like (ISO 18587:2017). So stop pretending your MT is done. Build a workflow that acknowledges the reality.
Scenario: You're a Localization Manager at a Mid-Sized Software Firm
Imagine you're the localization manager at a mid-sized software firm that's expanding into German, French, and Japanese. You've got a translation budget that's tight, a deadline that's tighter, and a CEO who read a blog post about DeepL's quality claims. He's convinced you can cut costs by feeding everything through Google Translate and shipping it. Your job is to convince him otherwise—and to build a process that works.
Step 1: Understand What MT Can and Cannot Do
First, get the facts straight. Neural machine translation (NMT) is indeed a leap forward: Google's GNMT, announced in September 2016, reduced translation errors by 55-85% on sampled Wikipedia and news sentences compared to the previous phrase-based system (GNMT at production scale, Google Research). But that doesn't mean it's perfect. NMT translates whole sentences, not word-by-word, but it still stumbles on context, tone, and brand voice (Machine translation, Wikipedia). And while DeepL claims its translations are chosen roughly three times more often than Google's in blind tests with professional translators (DeepL Translator, Wikipedia), that still leaves a lot of room for error. So, your first step is to set expectations: MT is a starting point, not a finish line.
Step 2: Choose Your Metrics Wisely
Now, how do you measure MT quality? If you're like most people, you've heard of BLEU. But BLEU is increasingly seen as unreliable because it correlates poorly with human judgment (IBM Research). So, don't rely on BLEU alone. Look at newer metrics like COMET, which achieved top performance at WMT 2019 and 2020 (COMET, ACL Anthology), or BLEURT, which correlates better with human judgments than BLEU (BLEURT, Google Research). For post-editing effort, consider Human-targeted Translation Edit Rate (HTER), which measures how many edits a human translator needs to make to fix MT output (Translation Edit Rate, ACL Anthology). These metrics give you a more realistic picture of how much work is left.
Step 3: Implement a Post-Editing Process with Clear Guidelines
Once you've measured, you need a process. The ISO 18587:2017 standard is your friend here. It specifies requirements for full human post-editing of MT output and for post-editors' competences (ISO 18587:2017). That means you need to define what 'done' looks like: Are you aiming for 'good enough' or 'publishable quality'? In your case, since it's software UI and help docs, you need near-publishable quality. So, you'll set up a workflow where MT output goes to a post-editor—a human translator who fixes errors, ensures consistency, and checks for cultural appropriateness. Don't skip this step. The MT output will have issues, and your post-editor is the safety net.
Step 4: Use CAT Tools and Translation Memory to Boost Efficiency
Here's where you can actually save money. Instead of translating everything from scratch, use a CAT tool with translation memory (TM). Translation memory is a database of previously translated segments that can be reused, ensuring consistency and reducing repeated work (Machine translation, Wikipedia). For example, if you've already translated a string like 'Save' or 'Cancel' in an earlier version, the TM will reuse it. This reduces the amount of new MT and post-editing needed. And when you do use MT, integrate it into your CAT tool as a fallback, just like Google Translator Toolkit did—it pre-translated segments from TM and fell back on MT when no match was found (Google Translator Toolkit, Wikipedia). That's a smart model.
Step 5: Measure, Iterate, and Scale
Once you have your process, you need to measure its effectiveness. Use HTER to track post-editing effort over time. If your MT provider improves, your HTER should drop. If it doesn't, consider adjusting your MT settings or switching providers. Also, keep an eye on your quality metrics with COMET or BLEURT. For example, on a typical software string, you might find that MT gets a BLEU score of 40, but after post-editing, the HTER is 30%—meaning your post-editor is changing about one-third of the words. That's a number you can track and report to your CEO. Over time, as your TM grows and your MT model improves, you can scale your process to more languages and more content.
Step 6: Concrete Example: A German Quick-Start Guide
Let's make this real. Imagine you're localizing a 1,000-word quick-start guide into German. You feed it through an NMT system like DeepL. The output is decent, but you spot issues: the tone is too formal for your brand, and a few idioms are mistranslated. For instance, 'hit the ground running' might come out as 'Schlag den Boden laufend'—which is nonsense. Your post-editor fixes these, and you measure the effort: HTER is 25%, meaning she had to edit about 250 words. That's not trivial. But because you used a CAT tool, the TM had already recycled 40% of the content from previous guides, so the actual new translation was only 600 words. The post-editor spent two hours on it. Without the MT and TM, it would have taken a full day. That's a real, measurable win.
Step 7: The Bigger Picture: Standards and Industry Reality
Finally, remember that you're not alone. The EU's Directorate-General for Translation, one of the largest translation services in the world, produced about 2.6 million translated pages in 2022, and it uses eTranslation, a free neural MT service, as an aid—not a replacement (European Commission translation department; eTranslation). WIPO Translate is another example: a neural MT tool for patents, but it's used to help humans, not replace them (WIPO Translate). And the language industry is big business: CSA Research estimated the global language services and technology industry at US$49.68 billion in 2023 (CSA Research). So, investing in a solid post-editing workflow is not a cost—it's a necessity for quality and consistency.
The One Thing to Remember
Machine translation is a tool, not a translator. Use it to speed up your workflow, but never ship MT output without human post-editing. Your brand, your users, and your sanity depend on it.
Sources
- Machine translation (Wikipedia) - https://en.wikipedia.org/wiki/Machine_translation
- ISO 18587:2017 (ISO) - https://www.iso.org/standard/62970.html
- GNMT at production scale (Google Research) - https://research.google/blog/a-neural-network-for-machine-translation-at-production-scale/
- COMET (ACL Anthology) - https://aclanthology.org/2020.emnlp-main.213/
- Translation Edit Rate (ACL Anthology) - https://aclanthology.org/2006.amta-papers.25/
- European Commission translation department - https://commission.europa.eu/about-european-commission/departments-and-executive-agencies/translation_en
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!