False positives

    Can Human-Written Work Be Flagged as AI? False Positives Explained

    Understand false positives, mixed fairness evidence and the process material that should be reviewed before any conclusion about a writer.

    Genutext Editorial Team4 min read
    Fairness evidenceBase-rate exampleReview workflow
    On this page
    1. What a false positive is
    2. Why human writing may look AI-like
    3. Formal and second-language writing
    4. Why low prevalence changes the risk
    5. How to review a disputed flag
    6. Frequently asked questions

    Yes. A false positive occurs when an AI detector labels genuinely human-written text as likely AI-generated. Every classifier has an error rate, and the risk is not evenly distributed across every text type, length or decision threshold.

    A false positive is especially serious in education because a statistical flag can be mistaken for evidence of misconduct. The safe response is to investigate the writing process—not to assume the model has recovered authorship.

    What a false positive is

    AI detection has four basic outcomes:

    Actual originDetector says humanDetector says AI
    Human-writtenTrue negativeFalse positive
    AI-generatedFalse negativeTrue positive

    Accuracy percentages can hide the distinction. For a misconduct process, the false-positive rate may matter more than overall accuracy because the harm falls on human writers who are wrongly flagged.

    False positives do not mean a product is fabricating a result. They mean the model found patterns associated with its AI class even though the true origin was human. The classifier does not know the history; it infers from the final text.

    Why human writing may look AI-like

    Potential contributors include:

    • a short sample with too little evidence;
    • formulaic academic phrasing or a rigid template;
    • repetitive sentence structure;
    • highly polished, uniform prose;
    • technical genres poorly represented in training data;
    • translation or language-learning patterns;
    • heavy conventional editing; and
    • text near the model's decision threshold.

    These are not reliable “signs of AI” by themselves. Many disciplines reward consistent terminology and formal structure. A methods section, legal summary or lab report may reasonably sound more regular than a personal essay.

    Try the text in context

    Run a free AI-writing signal check

    Paste 300 characters to 350 words. The sample is analysed for a first-pass signal and then discarded.

    The checker loads as you approach it.

    Formal and second-language writing

    The fairness evidence is mixed and product-specific.

    A widely cited 2023 study found that several detectors frequently misclassified essays by non-native English writers. The researchers linked this to lower linguistic variability in the tested material and showed why using detector output as proof could create unequal harm.

    Turnitin later reported testing nearly 2,000 English-language-learner samples and said its own false-positive rate was not statistically different from its native-English sample when documents met its 300-word requirement. That is relevant vendor evidence, but it does not erase the independent finding or guarantee equal performance for every language background, genre and model version.

    The defensible conclusion is narrower: do not assume that formal or second-language writing is AI-generated, and do not transfer a fairness claim from one detector to another. Review validation evidence for the exact tool and text type.

    Why low prevalence changes the risk

    Suppose 1,000 submissions include 50 genuinely AI-generated papers. Imagine a detector catches 90% of those and falsely flags 2% of human papers:

    • 45 AI papers are correctly flagged;
    • 19 human papers are falsely flagged;
    • 64 papers are flagged in total.

    In that example, almost three in ten flags are false even though the detector has strong sensitivity and a low false-positive rate. When the behaviour being detected is uncommon, false positives form a larger share of alerts.

    This base-rate effect is one reason institutions should not turn a percentage into an automatic sanction.

    How to review a disputed flag

    Check the report

    • Was the sample long enough and in a supported language?
    • Which passages were eligible and highlighted?
    • Is the score low or near a suppressed/uncertain range?
    • Which model version produced it?

    Check the writing process

    • outlines, notes and reading records;
    • document version history;
    • cited sources and quotations;
    • earlier work used only where policy permits and context is comparable;
    • the student's explanation of the argument and key choices; and
    • any declared or permitted support.

    Check the procedure

    The student should receive the allegation, the relevant report and the academic analysis, then have a meaningful chance to respond. The OIA's UK good-practice framework treats detection-software interpretation as evidence-based academic judgement, not an automatic fact.

    Read the complete educator workflow in what to do when a student disputes an AI result.

    Frequently asked questions

    Can 100% human writing be detected as AI?

    Yes. A high score can still be a false positive. The strength of a model output does not replace evidence about how the work was produced.

    Are non-native English writers always more likely to be flagged?

    No universal claim is supported. Some independent research found serious disparities across tested detectors, while Turnitin reports no significant difference in its own eligible sample. Risk depends on the tool, version, genre and population.

    Does formal vocabulary cause AI detection?

    No single word or phrase proves AI use. A classifier evaluates patterns across text, and formal vocabulary is normal in many academic genres.

    What evidence can show that work was human-written?

    Contemporaneous notes, drafts, version history, source records and the writer's ability to explain the argument are useful together. No single artefact is infallible.

    Sources and further reading

    Continue the topic

    Related Genutext guides

    View every guide

    Apply the guidance

    Run a first-pass check, then review the context

    Use the free AI preview for a short sample, or sign in for longer AI and plagiarism scans with sentence-level context.