2026: The Culture of AI Lost in Translation: An Autonomous AI and Language Translation Study
- WAI CONTENT TEAM

- 2 days ago
- 5 min read

Principal author: Karen Jensen
Welcome to the first published study authored and contributed by current and former team members, as we focus on the Culture of AI! In 2026, we continue our global initiatives in Education, Entrepreneurship, Innovation, and Research to make AI accessible and inclusive for everyone, with a special focus on women and girls.
This post introduces our study — Lost in Translation: An Autonomous AI and Language Translation Study — now available on ResearchGate. It’s an applied research white paper that evaluates six Large Language Models across two languages, Mandarin and Hindi, using native speaker review, a Forward and Reverse Translation Audit conducted by an IEEE CertifAIEd™ Lead Assessor, and a test we designed specifically for this study: asking the models to evaluate themselves.
The full white paper is linked throughout this post. What follows is the story behind it.
Where This Started
This study traces directly back to our 2025 Expert Series. In our seventh session, Intersectional Bias: Localization Challenges in AI and Solutions, Ishween Kaur made the case that language localization should not be a technical afterthought.
Despite over 63% of the world's population speaking a primary language other than English, Large Language Models (LLMs) are overwhelmingly trained on English text. This results in a one-dimensional practice that causes serious societal harm beyond simple translation failures.
That framing stayed with our study author, Karen Jensen. It also became the origin of this study.
As we considered translation challenges, we chose our source document deliberately: our final 2025 Expert Series blog post. That document contains technical terms, specific AI terms, and gender centric phrases, written for a global audience.
We then asked six Large Language Models to translate it, both forward and reverse, into Mandarin and into Hindi.
Below are the results.
What We Tested and How
Six models participated in the study: Gemini Flash 3, ChatGPT 5.3, 5.4, and 5.5, Claude Sonnet 4.6, Perplexity.ai, and Apertvs. These models represent both proprietary/commercial models and an open-source model (Apertvs).
We chose two languages deliberately. Mandarin is a high-resource language — it has vast training data. Hindi, spoken by more than 600 million people, is classified as a low-resource language. High and low language designations are not based on the number of speakers but on the digital content written in it that models use for training.
Our evaluation framework had three components:
• Native speaker review, using a five-dimension rubric covering Semantic Fidelity, Linguistic Accuracy and Sovereignty, Cultural Nuance and Resonance, Terminology Precision, and Localization Fluency.
• A Forward and Reverse Translation Audit conducted by Vanessa Maidoh, IEEE CertifAIEd™ Lead Assessor, under IEEE 7003-2024. Each model translated the source into the target language, then translated its own output back to English. The gap between those two steps is where bias is masked.
• A five-prompt autonomous evaluation sequence in which the models were asked to predict their own scores, evaluate the full set of translations blind, identify which translation was their own, and then reflect when the true answer was revealed.
What We Found
The short version: no model excelled across all dimensions. Every model had documented failures.
In Mandarin, the most consistent failure was terminology. “Agentic AI” became “proactive AI” in two models — a translation that loses the entire meaning of the term. “Explicit content” was softened or mis-rendered in two models. “Intersectional bias” was mistranslated as “intertwined bias” — a term that sounds plausible but describes something entirely different.
In Hindi, the pattern shifted. Nearly every model scored a 1 on Localization Fluency — meaning the output did not read as native Hindi. One model explicitly announced within the translated text that a translation was being performed. Another skipped an entire large section of the source document.
Bias Laundering: What English-Only Readers Cannot See
Vanessa’s audit used a methodology developed specifically for this study: comparing the forward translation against the reverse translation using an English-only auditor who does not read the target language.
This is precisely the situation most organizations face when they deploy AI for translation. They see the English output. They do not see what happened in between.
What the audit found: a male default in Mandarin in a document written entirely about women in AI. In the reverse translation, that pronoun became the gender-neutral “who.” An English-only reader would never know.
Four of five study models substituted “gender equality” for “gender parity” in the closing statement of the Women in AI mission. These are not synonyms. Gender parity is a specific, measurable policy term. The substitution was visible in the reverse translation but would not be obvious to a casual reader.
Bias Laundering — an established concept in algorithmic bias — describes exactly this: a model introduces an error or bias in the forward translation, then restores the appearance of neutrality in the reverse translation, masking the failure from anyone who only reads English.
We Asked the Models to Evaluate Themselves. None of Them Got It Right.
Four models completed a five-prompt autonomous evaluation sequence. We asked them to predict their own scores, evaluate the full translation set blind, identify which translation they believed was their own, and reflect when the true answer was revealed.
Every model overestimated its own performance. Every model identified a higher-scoring file as its own output. No model correctly recognized its own translation.
Why This Matters for Organizations
Organizations are deploying AI translation at scale.
The framework in this study — rubric-based native-speaker evaluation, forward- and reverse-translation audit, autonomous model evaluation — is reproducible. You do not need to conduct a six-model study. But you do need a framework for knowing whether the translation you are deploying can be trusted. This study gives you one.
The Human in the Loop Is Not Optional
Before Large Language Models, professional translation had a non-negotiable standard: ISO 17100:2015, which mandates human revision as a required step in any professional translation workflow. The American Translators Association has certified human translators since 1959.
LLMs did not change that requirement. They made it easier to overlook.
This study demonstrates that the failure modes in AI translation — gendered defaults, terminology substitutions, content omissions, and bias introduced in one language and masked in another — are not detectable without expert human review
Read the Full Study
Lost in Translation: An Autonomous AI and Language Translation Study is now available on ResearchGate.
The study is grateful for the contributions of:
Charlotte Tao: native Mandarin speaker and former Women in AI volunteer.
Ishween Kaur: native Hindi Speaker, and 2025 Women in AI Volunteer of the Year.
Vanessa Maidoh: IEEE CertifAIEd™ Lead Assessor and Women in AI volunteer
Karen Jensen: study author and Women in AI volunteer
The question this study started with — can AI translate our content accurately for a global audience — has a clear answer now. Not without a framework. Not without a human in the loop.
This blog post is part of the Women in AI Global Office of Ethics and Culture’s commitment to a Global Vision for achieving gender parity in emerging technologies through increasing Opportunity, championing inclusive Policies, and fostering practical Action that delivers meaningful and measurable impact.
Ethics & Culture Team

Please see the links below to our Team’s profiles on LinkedIn.



Comments