AI NewsProductivityReported

Human translators lose 4 of 6 English-to-Chinese content types to post-edited AI, and one model choice matters more than post-editing

A 774-output English-to-Chinese benchmark from EC Innovations and Jademond Digital, published in Search Engine Journal on 28 September, ranked professional human translation outside the top five on marketing, technical, product UI and user-generated content, while post-edited Qwen won or tied four categories.

AI News

Editorial3 min read

LinkedInX

Why it mattersA team routing localization spend by content type, and picking a specific Chinese model rather than the average of a category, saves more than a team debating whether to post-edit at all.

A working localization budget is being set the wrong way. A 774-output benchmark of English-to-Chinese translation ran professional human translators 10th of 15 on marketing copy, 22.2 points behind post-edited Qwen, and outside the top five on four of six content types.

Search Engine Journal published the results on 28 September, in a piece by Marcus Pentzek that draws on a joint study by EC Innovations and Jademond Digital, where Pentzek is a Partner and Director. Pentzek states on the page that both firms sell services related to the practices the research evaluates. The full methodology and report is on Jademond's site.

What was measured

Professional human localization was tested against fourteen machine and hybrid workflows across six content types: informational, SEO, technical, product UI, user-generated content and marketing. Each output was scored by Chinese native-speaking professional localizers, blind to which workflow produced it, on accuracy and consistency, fluency and language quality, and style and cultural adaptation, each dimension weighted equally at one third of the total.

Humans finished first on informational (76.9) and SEO (74.1). Their SEO lead over the next workflow, post-edited Doubao and post-edited Qwen tied at 71.3, was 2.8 points. Humans finished 7th on technical (64.8, behind PE-Qwen at 79.6), 7th on product UI (63.0, behind PE-Qwen at 73.1), 9th on user-generated content (64.8, behind PE-Doubao and PE-Qwen tied at 75.9), and 10th on marketing (53.7, behind PE-Qwen at 75.9).

Model choice mattered more than post-editing

Search Engine Journal reports that post-editing lift averaged 5.6 points across the study, while the gap between the best and worst Chinese model in a single content type ran up to 22.3 points. Post-editing three cases, PE-ChatGPT on marketing at minus 3.7, PE-Kimi on marketing at minus 2.8, and PE-ChatGPT on technical at minus 0.9, made the output worse. Kimi finished last of the four Chinese models in all six content types. Search Engine Journal notes that removing Kimi from the "Chinese LLM" category average, which any two-week enterprise trial would do, shifts the figure by 1.3 to 4.9 points.

The caveats the author states himself

Pentzek writes on the page that the study measured localization quality only, and collected no ranking, traffic or conversion data. He points to a 2021 Portent crawl of 756,297 pages that found no correlation between readability and Google ranking position, and calls out the circularity of research funded by content-quality vendors that then claims content quality drives rankings, including his own. He asks readers to treat every finding here as an input hypothesis to test on their own pages.

Two more caveats sit on the page. Every model in the study was tested at its December 2025 version, so Qwen, Doubao, GPT and Gemini have all shipped since. The LLM and MT systems tested were GPT-5.2 (via ChatGPT), Gemini 3.0, Doubao 1.6, Qwen 3, Kimi K2, DeepSeek-V3.2 and Google Translate, accessed through their web interfaces.

For a team spending on Chinese localization, the numbers move the question from "human or post-edit" to "which model, in which category, and at what stage in the site build". Chinese SEO carries a filing and hosting prerequisite that this study does not touch, and Pentzek writes that upgrading from post-edited Qwen to full human translation, on top of a broken ICP filing or slow China hosting, is the more expensive of the two fixes and buys the smaller improvement.

Source

Primary: Jademond Digital, English-to-Chinese Localization Benchmark Report. Reporting: Search Engine Journal, AI workflows outscored human translators in 4 of 6 content types in a China benchmark study.

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX