Skip to main content
Tools & Software

Why Your Translation Stack Needs a Human in the Loop

Think machine translation alone will do? It won't. Here's why CAT tools and post-editing are non-negotiable for quality translations that actually meet industry standards.

Should I use machine translation or a human translator? That's the question every project manager asks when a new localization request lands on their desk. The answer is simpler than the vendor pitches suggest: neither alone. The right stack combines machine translation (MT) with computer-assisted translation (CAT) tools and human post-editing. That's the only way to hit the quality bar that clients actually pay for.

The false choice between MT and CAT

MT is fully automated and replaces human translation during the translation phase. CAT is machine-assisted human translation where the translator keeps control. They are not competitors; they are sequential steps. A modern workflow uses MT to generate a first draft, then a human post-editor refines it. The translator isn't replaced; they're repositioned. CAT tools like SDL Trados, memoQ, or OmegaT store previous translations in a translation memory (TM) database, ensuring consistency and reducing repeated work. When a segment matches the TM, it's reused; when it doesn't, MT fills the gap. That's exactly how Google Translator Toolkit worked from 2009 to 2019: it split documents into segments, pre-translated from TM, and fell back on MT when no match was found. The toolkit is gone, but the principle is standard practice today.

Post-editing is not optional

MT output often requires human post-editing, especially for idiomatic expressions, cultural nuance, and domain-specific terminology. Skipping that step is how you end up with embarrassing errors that cost more to fix than the translation itself. ISO 18587:2017 specifies requirements for full human post-editing of MT output and for post-editors' competences. If your vendor doesn't follow this standard, you're gambling. And don't confuse post-editing with proofreading: it's a specialized skill. The industry even has a metric for it: Human-targeted Translation Edit Rate (HTER) measures the edits a human must make to bring MT output up to publishable quality. Use it to estimate effort.

Evaluation metrics matter more than ever

BLEU, introduced by IBM researchers in 2002, scores translations on a 0 to 1 scale by comparing n-gram matches. But BLEU correlates poorly with human judgment. That's why newer metrics like COMET and BLEURT are becoming standard. COMET, developed by Unbabel, won the WMT 2019 and 2020 metrics shared tasks. BLEURT, from Google Research, uses BERT and correlates better with human judgments on several generation tasks. For morphologically rich languages, chrF is a better choice because it's character-based and tokenization-independent. Don't rely on a single number; use a combination. And remember: metrics guide development, but they don't replace human review.

The counter-argument: neural MT is good enough

Some claim that neural MT has improved so much that post-editing is unnecessary. In September 2016, Google announced GNMT, which reduced translation errors by more than 55–85% on several major language pairs compared with the previous phrase-based system. DeepL, launched in 2017, reported that in blind tests its translations were chosen roughly three times more often than those of Google, Microsoft, or Facebook. Those are impressive gains. But 'good enough' depends on the use case. For a legal contract or a medical label, a 10% error rate is unacceptable. For internal emails, it might be fine. The mistake is treating MT as a one-size-fits-all solution. You still need a human to judge when MT output crosses the line.

Quick tip: Always run a sample through your MT engine and have a human post-editor review it before committing to a full project. That 30-minute test can save you thousands in rework.

Build your stack around standards

Use XLIFF for exchanging bilingual files. XLIFF 2.1 was approved as an OASIS Standard on 13 February 2018, and it's supported by major CAT tools. Microsoft lists XLIFF, TBX, and TMX as the standardized interchange formats for CAT tools. These aren't just technical details; they're what allow you to switch vendors without losing your translation memory. And if you're buying translation services, demand ISO 17100:2015 compliance. It specifies requirements for core processes, resources, and other aspects of delivering a quality translation service. If a vendor can't show you their ISO 17100 certificate, walk away.

The single most important thing to remember: machine translation is a tool, not a strategy. Pair it with CAT tools, enforce post-editing standards like ISO 18587, and measure with modern metrics. That's how you deliver translations that pass scrutiny.

Sources

  • Machine translation (Wikipedia) - https://en.wikipedia.org/wiki/Machine_translation
  • ISO 18587:2017 (ISO) - https://www.iso.org/standard/62970.html
  • IBM Research - https://research.ibm.com/blog/bleu-nlp-benchmark-anniversary
  • Google Research (GNMT) - https://research.google/blog/a-neural-network-for-machine-translation-at-production-scale/
  • XLIFF 2.1 (OASIS) - https://docs.oasis-open.org/xliff/xliff-core/v2.1/xliff-core-v2.1.html

Share this article:

Comments (0)

No comments yet. Be the first to comment!