Methodology

Do AI Detectors Work on Non-English and Translated Text?

Validation is per language, a translation is a different document, and machine-translated prose is itself generated text — what that means for scores, students and reviewers.

Genutext Editorial Team7 min read
3 languages Turnitin scoresTranslation as a special case5 FAQs
On this page
  1. The short answer
  2. Why language changes the score
  3. Which languages the major tools actually score
  4. Translated text: the special case
  5. What this means for students and reviewers
  6. Frequently asked questions

Mostly, no — or at least not the way they work on English. AI-writing detectors are statistical models trained and validated overwhelmingly on English prose, so the same ideas expressed in another language are a different statistical object, scored by a model that may never have been tested on that language at all. And text that has been machine-translated occupies the worst position of all: translation software is itself a neural text generator, so its output can carry exactly the patterns detectors are trained to flag.

The short answer

The same text does not get "the same score in a different language" — it gets a different score, from a model working outside the conditions it was validated for. Three things drive that:

  • Detectors are validated per language. A detector's published accuracy figures come from test sets, and those test sets are overwhelmingly English. Turnitin, the most consequential tool in education, scores qualifying prose only in the languages it lists — its AI writing report documentation names English, Spanish and Japanese — and text outside its supported languages simply is not scored.
  • A translation is a different document. Sentence length, word frequency, burstiness and phrasing — everything a detector measures — change when the language changes. Comparing a 34% on the English version with a 12% on the Spanish version tells you about the models, not the text. This is the cross-language version of a problem we cover in depth in why AI detectors give different scores.
  • Unsupported does not mean safe. A tool asked to score a language it was not built for may refuse, silently score only fragments, or produce a number with no validation behind it. None of those outcomes is evidence about authorship.

Try the text in context

Run a free AI-writing signal check

Paste 300 characters to 350 words. The sample is analysed for a first-pass signal and then discarded.

The checker loads as you approach it.

Why language changes the score

Detectors estimate how "machine-like" prose is: how predictable each word is given the last ones, how uniform the sentence rhythms are, how the phrasing compares with the model's training distribution. All of those measurements are anchored to a language's own statistics.

Move the same argument from English into French, Yoruba or Mandarin and every anchor moves. If the detector's training data for that language is thin, ordinary native prose can look "unusual" to it — or eerily regular, which reads as machine-like. Independent testing bears out how fragile the results are even inside English: the Weber-Wulff multi-tool study found no tool it tested exceeded 80% accuracy, and accuracy fell further on manipulated text. Outside English, most tools publish no equivalent evidence at all — which means a score on non-English text is a number without a stated error rate.

There is a well-documented adjacent finding: Stanford researchers (Liang et al., 2023) showed that essays written in English by non-native speakers were flagged as AI at far higher rates than native-speaker essays — seven detectors misclassified more than half of a TOEFL essay set. That study is about writers, not languages, and we cover it separately in AI detectors and non-native English writers; but it demonstrates the underlying mechanism, because simpler, more uniform phrasing — whether from a learner or a translation engine — is precisely what detectors read as machine-generated.

Which languages the major tools actually score

The honest answer for any tool is "check its documentation on the day you use it" — language support changes, and a vendor's marketing page and its technical limits are not always the same list. Two useful, stable anchors:

  • Turnitin documents its AI writing report as scoring qualifying long-form prose in English, Spanish and Japanese, with a 300-word minimum; content outside the supported languages is excluded from the percentage. So a submission in German or Arabic does not produce a meaningful Turnitin AI score at all — a point worth knowing before anyone reads a number off a report. Our guide to reading the Turnitin AI report covers what else the indicator excludes.
  • Genutext scans English prose and says so: the accuracy methodology sets out the supported input, and text outside it should not be scored as if the validation applied.

For other tools, apply the same two questions you would ask of any score — what languages was this validated on? and what happens to unsupported text? If the documentation doesn't answer, treat the score accordingly.

Translated text: the special case

Text translated by DeepL, Google Translate or an LLM sits in a category of its own, because the words on the page were literally produced by a neural text generator — whatever the ideas' origin.

  • Machine-translated output can score as AI. Translation engines produce fluent, statistically regular prose; regularity is the core AI signal. A student who writes in their first language and machine-translates into English has human ideas wrapped in machine-generated sentences, and detectors cannot see the difference between that and generated content.
  • Policies increasingly treat translation as AI assistance. Several UK universities' generative-AI guidance addresses translation tools explicitly — some permitting them with declaration, some not — and our UK policy tracker links each institution's own statement. Goldsmiths, for instance, cites unfairness to students who used AI-powered translation as one reason it has not enabled detection.
  • Self-translation is safer than machine translation for detection risk — prose you compose yourself in English carries your own irregularities — but the real protection is process evidence: drafts in the original language, translation notes, version history.

What this means for students and reviewers

  • Students: if you write in another language or translate your own work, keep the original-language drafts. If your institution permits translation tools, declare them where a declaration exists. If you're facing a score-based allegation, our guide to proving you didn't use AI covers the evidence that carries weight.
  • Reviewers: before reading any percentage on non-English or possibly-translated work, check the tool's supported-language list. A score on unsupported text has no stated error rate, and UK guidance is consistent that a detector score alone is not evidence of misconduct in any language.

Frequently asked questions

Does an AI detector give a different score for the same text in different languages?

Yes, almost always. Each language version is a statistically different document, and most detectors are trained and validated mainly on English — so the scores differ, and neither number transfers to the other language. Only the version in a language the tool documents support for carries the tool's published accuracy.

Can Turnitin detect AI in languages other than English?

Turnitin's documentation lists English, Spanish and Japanese for its AI writing report. Text in other languages is not qualifying content for the AI percentage, so a report on, say, a German essay is not scoring the German prose.

Does Google Translate or DeepL text get flagged as AI?

It can. Machine translation is neural text generation, and its fluent, regular output resembles what detectors are trained to flag — even when the underlying ideas and structure are entirely the writer's own.

Is it safe to write in my language and translate my essay into English?

Check your university's generative-AI policy first — institutions differ on whether translation tools are permitted or must be declared. For detection risk specifically, machine-translated text is more likely to be flagged than prose you composed in English yourself, so keep your original-language drafts as evidence either way.

Are there AI detectors built for other languages?

Some vendors advertise multi-language support, but published validation outside English is scarce. Ask for the language-specific accuracy evidence; if none is published, treat the score as unvalidated.

Sources and further reading

Continue the topic

Related Genutext guides

View every guide

Writing in English as a second language?

See which sentences read as machine-generated

A sentence-level Genutext scan shows which passages of your English draft carry AI-like signals — useful before you submit translated or self-translated work. Your first account scan is a free preview.