False positives

    AI Detectors and Non-Native English Writers: What the UK Evidence Shows

    The Liang and HEPI evidence UK-side, the mechanism behind the bias, and practical protection for international students.

    Genutext Editorial Team6 min read
    61% misclassified (HEPI)Peer-reviewed basis5 FAQs
    On this page
    1. The short answer
    2. The founding evidence
    3. The UK picture in 2026
    4. Why detectors penalise non-native writing
    5. Does language or translation change a score?
    6. What international students can do
    7. What vendors say
    8. Frequently asked questions
    9. Sources and further reading

    If English is not your first language, your honest writing is statistically more likely to be flagged as AI. That is not folklore — it is one of the best-documented failure modes in detection research, and in 2026 it became a UK policy issue. Here is the evidence, why it happens, and what you can practically do.

    The short answer

    Peer-reviewed research found public GPT detectors misclassified a majority of non-native English writing samples while staying near-perfect on native samples; UK sector analysis in 2026 reported 61% of non-native speaker essays misclassified in the research it reviewed, and framed the risk around the UK's international students — about a quarter of the student body and half of postgraduates. If you are an international student, this cuts both ways: you are at higher false-positive risk, and that risk is now well-enough documented that citing it in a dispute is legitimate, not special pleading.

    The founding evidence

    The study everyone cites is Liang et al. (2023), published in Patterns: GPT detectors are biased against non-native English writers. Testing widely-used public detectors on TOEFL essays written by non-native speakers versus essays by native speakers, the authors found the non-native essays were misclassified as AI-generated at dramatically higher rates — and that simple prompts to "enhance" the language flipped classifications, showing the detectors were keying on linguistic simplicity rather than actual machine origin.

    Subsequent cross-tool testing (Weber-Wulff et al., 2023, IJEI) reinforced the broader reliability picture: accuracy varies by tool and text type, and confident classifications of human writing as AI occur at rates no misconduct process should ignore.

    The UK picture in 2026

    Three UK-specific data points, each linked:

    • HEPI (20 July 2026) — the Higher Education Policy Institute's analysis reports 61% of non-native English essays misclassified in the research it reviews, sets it against detection accuracy of 39.5% on unmodified AI text (17.4% after simple edits), and notes international students are ~24% of UK higher education and 51% of postgraduates. It calls for sector guidance prohibiting detection scores as a sole disciplinary basis.
    • Times Higher Education (23 February 2026) — a YouGov/Studiosity survey of 2,373 UK students found 75% of AI-using students stressed about being wrongly flagged, with 52% citing fear of being accused of cheating they didn't commit.
    • University behaviour. Several UK universities' published reasons for not running Turnitin's AI indicator cite fairness and reliability; student-press reporting has highlighted false accusations of neurodivergent and international students.

    Why detectors penalise non-native writing

    Detectors estimate how predictable prose is — machine text tends to be statistically smooth. Non-native academic English often shares that smoothness for innocent reasons:

    • Learned formality. Writers taught academic English from templates produce regular, conservative sentence structures.
    • Smaller active vocabulary lowers lexical surprise — the exact signal detectors read as machine-like (Liang et al.'s mechanism).
    • Grammar and translation tools further regularise phrasing; some detectors also class heavy tool-polishing as AI-influenced.
    • Discipline conventions (methods sections, report formats) compound the effect regardless of language background.

    None of these is misconduct. All of them raise scores. That gap is the false-positive problem, and it is not unique to non-native writers — they just experience it most.

    Try the text in context

    Run a free AI-writing signal check

    Paste 300 characters to 350 words. The sample is analysed for a first-pass signal and then discarded.

    The checker loads as you approach it.

    Does language or translation change a score?

    Yes, in ways worth knowing:

    • Same text, different language: most detectors are trained mainly on English; many support only English or a short language list (Genutext's preview is English-only). Scores on non-English text are less validated and can differ sharply between tools.
    • Translated text: machine-translating your own work into English typically raises AI scores, because the translator's output is machine-generated prose in the statistical sense — even though the ideas are yours. If you draft in your first language and translate, keep both versions as process evidence, and disclose translation tools where your module requires it.
    • Polished text: running text through heavy paraphrase or "enhancement" can move scores in either direction (Liang et al. showed both flips). Polishing to evade a detector is the one move that converts an innocent situation into a policy breach — rewrite in your own voice instead.

    What international students can do

    1. Keep everything. Drafts in your first language, outlines, version history, dictionary/translation lookups — your process trail is the counter-evidence that decides disputes.
    2. Know your module's rules on translation and grammar tools, and disclose what you used where disclosure is invited.
    3. Understand your own text's signals. A sentence-level check shows which passages read as machine-smooth so you can vary them deliberately — improving your writing's voice, not gaming a score.
    4. If flagged, cite the evidence calmly. The Liang and HEPI findings belong in your response — alongside, never instead of, your process evidence. Our step-by-step guide for accused students includes template wording.

    What vendors say

    Turnitin has published research reporting no statistically significant bias against English language learners for its own tool — a claim specific to its model and benchmark, sitting alongside the independent findings above about public detectors. The fair reading: bias is demonstrated for public tools, contested for Turnitin, and consequential enough either way that UK guidance says no score should stand alone. Turnitin's overall accuracy evidence is reviewed here.

    Frequently asked questions

    Are AI detectors biased against non-native English speakers?

    Peer-reviewed research says yes for widely-used public detectors, with a majority of non-native TOEFL essays misclassified. Turnitin disputes it for its own tool. UK sector analysis treats the risk as real and current.

    Does an AI detector give a different score for the same text in different languages?

    Often, yes. Most detectors are optimised for English; support and validation for other languages varies by tool, so the same ideas can score differently across languages and tools.

    Will translating my own work make it look like AI?

    It can — machine translation produces statistically smooth prose that detectors may flag. Keep your original-language draft as evidence and disclose translation tools where required.

    I'm an international student who was flagged. What should I do?

    Preserve your drafts and version history, contact your students' union adviser, and respond with your process evidence plus the documented false-positive research. Follow the step-by-step guide linked above.

    How can I lower my false-positive risk without cheating?

    Vary sentence structure, use concrete personal examples, cite and quote properly, and keep process evidence. Do not use "humanisers" — they create the problem they claim to solve.

    Sources and further reading

    Genutext is independent of every organisation named here. Figures are quoted from the linked sources as of 23 August 2026.

    Continue the topic

    Related Genutext guides

    View every guide

    Apply the guidance

    Run a first-pass check, then review the context

    Use the free AI preview for a short sample, or create a free account — your first full scan (up to 3,000 words) is included.