One Number That Says It All
In 2016, Google announced that its neural machine translation system cut translation errors by 55–85% on several major language pairs compared with its previous phrase-based system (GNMT at production scale). That’s a leap that changed the industry. If you’re a translation buyer, that stat might tempt you to fire your human translators. Don’t. Not yet.
That error reduction is real, but it’s not the whole story. Machine translation (MT) still stumbles on idioms, cultural nuance, and domain terminology (Machine translation). And the language services market is shrinking—CSA Research pegged it at $49.68 billion in 2023, down 4.5% from the year before (CSA Research). That squeeze means you can’t afford to waste money on the wrong approach.
So here’s the fight: pure human translation, raw machine translation, or a hybrid—machine pre-translation plus human post-editing. I’m going to compare them on cost, quality, speed, and scalability. Then I’ll tell you which one wins, and under what conditions.
Human Translation: The Gold Standard, But Slow and Pricey
Human translation is the benchmark. A professional translator doesn’t just swap words; they capture tone, intent, and cultural context. That’s why ISO 17100:2015 exists—it sets requirements for the core processes and resources needed to deliver a quality translation service (ISO 17100). It’s a serious discipline.
But human translation is expensive and slow. It doesn’t scale. If you need to translate a 10,000-page patent portfolio into 10 languages, you’re looking at months of work and a six-figure bill. And even then, humans make mistakes.
Yet for high-stakes content—legal contracts, medical instructions, marketing copy that will make or break a brand—human translation is non-negotiable. That’s not nostalgia; it’s risk management. A single mistranslated clause can cost millions in litigation.
Who is this for? Companies with deep pockets, low volume, and zero tolerance for error. Think pharmaceutical labeling, where a mistranslated dosage could kill someone.
Raw Machine Translation: Fast, Cheap, and Often Wrong
Machine translation has come a long way since the Georgetown-IBM experiment in 1954, which used a 250-word vocabulary and six grammar rules (Georgetown-IBM). Today’s neural MT, like DeepL, Google Translate, and Microsoft Translator, uses deep learning to analyze whole sentences, not word-by-word (Machine translation). The Transformer architecture, introduced in 2017, made this possible (Attention Is All You Need).
The upside? Speed and cost. Raw MT can translate millions of words in seconds for pennies. For gisting—understanding the gist of a foreign-language document—it’s unbeatable. For low-stakes content like internal emails or user-generated reviews, it might be good enough.
The downside? Quality is unpredictable. BLEU, the classic automated metric, correlates poorly with human judgment (IBM Research). Newer metrics like COMET and BLEURT are better, but they’re still not human. And even the best MT requires post-editing for idiomatic expressions, cultural nuance, and domain-specific terminology (Machine translation).
Who is this for? Tech-savvy individuals or startups with tight budgets and no legal exposure. If you’re translating a blog post for fun, raw MT is fine. But if you’re translating a user agreement that will be enforced in court, you’re playing with fire.
Hybrid: Post-Editing Machine Output
The hybrid approach is what the industry is actually moving toward. You run MT first, then have a human post-edit the output. This is so common that ISO 18587:2017 sets requirements for full human post-editing of MT output and for the post-editors’ competences (ISO 18587).
The key metric here is Human-targeted Translation Edit Rate (HTER), which measures how much a human has to fix the machine output (Translation Edit Rate). Lower HTER means less work, so you can estimate costs.
Hybrid workflows shine when you have high volume and a controlled domain. For example, the European Commission’s Directorate-General for Translation produced about 2.6 million pages in 2022 (European Commission). They use MT like eTranslation to handle the volume, then have human translators polish the output (eTranslation).
Who is this for? Mid-to-large enterprises with steady translation needs, like software companies localizing their UI, or law firms translating routine contracts. You get decent quality at a fraction of the cost of pure human translation, and you keep control.
Head-to-Head: Cost, Quality, Speed, Scalability
Let’s put them side by side. Cost: raw MT wins—it’s nearly free. Human translation is the most expensive, often 10–20 times more per word than MT. Hybrid sits in between: you pay for the post-editor’s time, but you save on the initial draft.
Quality: human wins. A professional translator will catch nuances that MT misses. Hybrid is close behind, especially if you use a post-editor who’s trained in the domain. Raw MT is the worst, but it’s improving—DeepL claims its output is chosen three times more often than Google’s or Microsoft’s in blind tests (DeepL).
Speed: raw MT is instantaneous. Hybrid is faster than human-only, because the MT draft gives the post-editor a head start. Human translation is the slowest, especially for large projects.
Scalability: raw MT and hybrid scale to any volume. Human translation hits a wall when you need 100,000 words in a week—you’d have to hire an army.
| Criterion | Human | Raw MT | Hybrid (MT + Post-Edit) |
|---|---|---|---|
| Cost | High | Low | Medium |
| Quality | Excellent | Variable, often poor | Good–Excellent |
| Speed | Slow | Instant | Fast |
| Scalability | Limited | Unlimited | High |
The Verdict: Hybrid Wins—But Only With the Right Setup
So who wins? For most businesses, it’s the hybrid. It’s the only option that balances cost, quality, and speed. But you can’t just paste MT output and call it a day. You need to use tools that leverage translation memory (TM)—a database of previously translated segments that boosts consistency and cuts costs (Machine translation). CAT tools like Trados or memoQ integrate MT and TM, letting translators reuse past work.
And you must measure quality properly. Don’t rely on BLEU alone; it’s increasingly seen as unreliable (IBM Research). Use COMET or BLEURT, which correlate better with human judgment (COMET, BLEURT). Also, consider TER, which counts the edits needed to fix MT output—that’s a good proxy for post-editing effort.
But here’s the catch: hybrid only works if you have skilled post-editors. ISO 18587 spells out their competences (ISO 18587). If you don’t have them, you’ll get garbage. So invest in training or hire a language service provider that follows these standards.
One concrete example: suppose you’re a software company localizing your app into French, German, and Japanese. Your UI strings are short, repetitive, and high-volume—perfect for MT. You run them through DeepL, then have a native-speaking post-editor fix the output. The post-editor uses TM to ensure consistent terminology across versions. Result: you ship in three languages in a week, at half the cost of human-only, with quality that passes QA.
Now, if you’re translating a marketing slogan that will define your brand, don’t use MT. Even post-edited, the creativity may be lost. That’s when you call a human translator who understands the culture.
So my recommendation is not one-size-fits-all. It’s a decision tree: if you need gist, use raw MT. If you need perfection, use human. If you need volume and quality—which is almost everyone—use hybrid.
What I’d Actually Do
I’d build a workflow that treats MT as a first draft, not a final answer. I’d set up a CAT tool with translation memory, connect it to a high-quality NMT engine like DeepL or Azure AI Translator (which lets you customize with your own TM). I’d use COMET to spot-check quality, not BLEU. And I’d only hire post-editors who meet ISO 18587 standards.
If I were advising a client with a $100,000 budget for a 500,000-word project, I’d allocate 80% to hybrid post-editing and 20% to human review of the most critical sections. That’s the sweet spot.
Raw MT alone is a gamble. Human-only is a luxury. Hybrid is the workhorse—and it’s where the industry is heading. The numbers back it up: the language services market is declining, so you need to squeeze every dollar. Hybrid gives you the best ROI.
Don’t be fooled by the error-reduction hype. That 55–85% improvement is impressive, but it still leaves a 15–45% error rate. That’s too high for legal, medical, or financial content. For everything else, MT plus human judgment is the way.
Sources
- Machine translation (Wikipedia) - https://en.wikipedia.org/wiki/Machine_translation
- IBM Research - https://research.ibm.com/blog/bleu-nlp-benchmark-anniversary
- ISO 17100:2015 - https://www.iso.org/standard/59149.html
- ISO 18587:2017 - https://www.iso.org/standard/62970.html
- COMET - https://aclanthology.org/2020.emnlp-main.213/
- CSA Research - https://csa-research.com/l/media/Language-Services-and-Technology-Industry-Faces-Revenue-Decline-but-Remains-Poised-for-Transformation
Comments (0)
Please sign in to post a comment.
Don't have an account? Create one
No comments yet. Be the first to comment!