I keep coming back to one number: 55–85%. That's how much Google said its neural machine translation (GNMT) cut translation errors on sampled Wikipedia and news sentences in September 2016, compared with its old phrase-based system (GNMT at production scale, Google Research). It's a staggering improvement. But here's what that number doesn't tell you: it doesn't mean 55–85% of your localization work just disappeared. It means the raw material got better. The real question — the one I want to answer here — is whether that leap makes human translators obsolete for localization. My answer is an unqualified no, but with a big caveat: the role changes, and if you don't change with it, you're wasting money.
The question that actually matters
When people ask me if machine translation (MT) has killed localization, I reframe it: can a fully automated pipeline produce publishable, brand-safe, legally sound localized content across every language you care about? Because that's what localization is — not just swapping words, but adapting a product, a message, a legal disclaimer, a UI string, a marketing hook so it lands correctly in another culture. MT is a tool for the translation phase. Localization is the whole operation. Conflating the two is how companies end up with embarrassing product names and support tickets in languages they can't read.
So the narrow question is: where exactly does the human still earn their keep, and where is paying for one just inertia? I'll walk through the evidence, then tell you what I'd actually do.
What the neural leap did — and didn't — solve
Neural MT analyzes whole sentences rather than translating word by word, and it's trained end-to-end on large parallel corpora. That's a fundamental shift from the statistical era, when phrase-based systems dominated commercial MT before around 2016. Google's own multilingual NMT system, announced in November 2016, even enabled zero-shot translation — Japanese to Korean, for example — without ever being trained on that pair. Researchers found evidence the system built an internal interlingua, a shared representation across languages. That's genuinely impressive.
But impressive on Wikipedia and news is not the same as impressive on your 40,000-word software string file, your clinical trial consent form, or your luxury brand tagline. NMT output still stumbles on idiomatic expressions, cultural nuance, and domain-specific terminology. I've seen MT render a friendly product onboarding string into something that read like a bureaucrat's threat in German. The words were right. The tone was a lawsuit waiting to happen.
And the quality claims? DeepL, launched in August 2017, has reported that in blind tests with professional translators its output was chosen roughly three times more often than Google, Microsoft, or Facebook. That's a real signal — DeepL is good. But "chosen more often than competitors" is not "chosen over a human." Those are different tests. Don't let a vendor's comparative study become your excuse to skip human review.
Where humans still win — and where they don't
I split localization work into three buckets. Bucket one: high-volume, low-risk, repetitive content — internal knowledge base articles, user-generated support replies, non-legal product descriptions. Here, MT plus light post-editing is not just acceptable; it's the only sane economic choice. Bucket two: regulated or high-stakes content — legal, medical, financial, safety instructions. Here, MT output must go through full human post-editing. ISO 18587:2017 exists precisely to define what that looks like, including the competences a post-editor needs. If you're not following it, you're guessing. Bucket three: creative and brand-critical content — slogans, campaign copy, app store descriptions. Here, MT is a drafting aid at best. A human translator who understands the brand voice writes it, possibly after seeing an MT draft.
The measurement crowd will tell you to track post-editing effort with Human-targeted Translation Edit Rate (HTER), where a human creates a reference closest to the machine output and you count the edits needed. That's a useful proxy. But don't confuse a low HTER with a good localization. A string can need zero edits and still be culturally tone-deaf. Metrics measure distance from a reference, not fitness for a market.
The standards and formats that make this workable
What makes the hybrid model practical is the plumbing. XLIFF is the XML-based bilingual format standardized by OASIS, with XLIFF 2.1 approved as an OASIS Standard on 13 February 2018. Microsoft's localization guidance notes XLIFF 2.0 was adopted as ISO 21720:2017 and lists XLIFF, TermBase eXchange, and Translation Memory eXchange as the standardized interchange formats CAT tools use. Translation memory — a database of previously translated segments — lets you reuse approved translations so you're not paying a human to retranslate the same button label 12 times.
This is where computer-assisted translation (CAT) tools earn their place. CAT aids the human translator; it doesn't replace them. The old Google Translator Toolkit, which ran from 2009 to 2019, showed the model: split documents into segments, pre-translate from translation memory, fall back on MT when no match exists. That architecture is still the right one. It just runs on better MT now.
Quick tip: If your CAT tool isn't feeding MT output into empty segments and tracking what the post-editor actually changed, you're flying blind on cost and quality.
The economics — and why the market numbers are messy
Here's the part that makes CFOs nervous. CSA Research estimated the global language services and technology industry at US$49.68 billion in 2023, down 4.5% from US$52.01 billion in 2022. Slator's 2024 report put the same industry at US$27.03 billion in 2023. Those numbers are miles apart, and the reason is scope: whether interpreting and language technology are included. So when someone tells you "the translation market is shrinking because of AI," ask which definition they're using. A 4.5% dip in one estimate is not the same as an industry collapse.
What I do trust is the direction of travel: MT gets cheaper and better every year, and that pushes the value of human work up the stack. DeepL's EUR 1 billion valuation after its January 2023 funding round tells you investors believe machine translation is a huge business. It says nothing about whether your German legal contract can skip a human. It can't.
What I'd actually do
If you run localization for a company of any size, stop treating "MT vs. human" as a binary. Build a tiered pipeline. Route low-risk, high-volume content through MT with a light post-edit pass and a sampling QA process. Route regulated and brand-critical content through full human translation, using ISO 17100:2015 for the core service requirements and ISO 18587:2017 for post-editing. Invest in translation memory and XLIFF-based workflows so every human decision compounds. And measure post-editing effort honestly, not just with a single number like HTER but with a human review of tone and cultural fit.
The 55–85% error reduction is real. It means your baseline is better. It does not mean you can fire your translators. It means you can finally afford to use them where they matter most — and automate the rest. That's not a compromise. That's the whole job now.
Sources
- GNMT at production scale (Google Research) - https://research.google/blog/a-neural-network-for-machine-translation-at-production-scale/
- ISO 18587:2017 (ISO) - https://www.iso.org/standard/62970.html
- CSA Research - https://csa-research.com/l/media/Language-Services-and-Technology-Industry-Faces-Revenue-Decline-but-Remains-Poised-for-Transformation
- DeepL Translator (Wikipedia) - https://en.wikipedia.org/wiki/DeepL_Translator
- XLIFF 2.1 (OASIS) - https://docs.oasis-open.org/xliff/xliff-core/v2.1/xliff-core-v2.1.html
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!