Score interpretation

    What Does an AI-Detection Percentage Actually Mean?

    Distinguish text coverage, probability, confidence and product scores before translating a number into a real-world decision.

    Genutext Editorial Team4 min read
    4 score meanings5 reading questionsTurnitin example
    On this page
    1. Four different meanings of a percentage
    2. Coverage versus probability
    3. Confidence is not certainty
    4. How Turnitin defines its percentage
    5. A five-question reading method
    6. Frequently asked questions

    An AI-detection percentage does not have one universal meaning. Depending on the product, it may describe the share of eligible text highlighted, a model's confidence, a calibrated probability, or a score mapped to a label.

    Before interpreting the number, ask: percentage of what?

    Four different meanings of a percentage

    1. Text coverage

    The number may be the proportion of eligible prose classified as likely AI-generated. A 30% result then means about 30% of the qualifying text was highlighted—not a 30% chance of misconduct.

    2. Document probability

    A tool may estimate the probability that a document belongs to an AI-generated class. This is a model claim about a class under its calibration assumptions.

    3. Model confidence

    Some systems expose an internal confidence score. Confidence can be high and still wrong, especially when the text differs from the model's validation data.

    4. A normalised product score

    A provider may combine passage outputs into a 0–100 scale and map ranges to “human”, “mixed” or “AI”. Unless the methodology says it is a probability, do not read it as one.

    Try the text in context

    Run a free AI-writing signal check

    Paste 300 characters to 350 words. The sample is analysed for a first-pass signal and then discarded.

    The checker loads as you approach it.

    Coverage versus probability

    Consider a 2,000-word document containing 1,600 words of qualifying prose. If a detector highlights 400 eligible words, it might report 25% because 400 is one quarter of 1,600.

    That number does not mean:

    • there is a 25% chance the student used AI;
    • 25% of the entire file was written by ChatGPT;
    • the writer is 25% responsible for misconduct; or
    • the evidence satisfies a disciplinary threshold.

    The denominator matters because tools may exclude references, quotations, tables, code and short-form material. Compare highlighted passages with the exact report definition.

    Confidence is not certainty

    In machine learning, a confident output can still be a false positive. Calibration tells you how often similar confidence values are correct across a suitable test set; it does not reveal the origin of a particular document.

    Suppose a model is well calibrated and gives many documents a 90% probability. Across comparable documents, roughly nine in ten may belong to the predicted class. The remaining one in ten can still be wrong—and performance may shift for a new language, genre, model or editing process.

    This is why benchmark context matters more than a polished dial. Read how accurate AI detectors are for precision, recall and base-rate examples.

    How Turnitin defines its percentage

    Turnitin's March 2026 guide defines the overall AI percentage as the proportion of qualifying long-form prose its model determines could be AI-generated or AI-generated and then modified by paraphrasing or bypass tools.

    Turnitin also:

    • requires at least 300 words of prose;
    • supports specified languages and file types;
    • treats the AI percentage as independent of the Similarity score; and
    • suppresses numeric values below 20% because false positives are more frequent in that range.

    The 20% display rule is a product safeguard, not a universal line between human and AI writing. It should never be reused as a misconduct threshold.

    A five-question reading method

    1. What is the unit? Coverage, probability, confidence or product score?
    2. What text was eligible? Whole file, prose only, sentences or passages?
    3. What is the threshold? Which score becomes a label, highlight or alert?
    4. What validation applies? Same language, genre, length and model version?
    5. What decision is being considered? Screening, conversation or formal action?

    Then read the highlighted text and limitations. If a provider does not explain its score, the number should carry less weight.

    Genutext labels its output as an AI-writing signal and directs users to methodology and limitations. It does not describe the score as proof of authorship.

    Frequently asked questions

    Does 80% AI mean an 80% chance the text is AI-written?

    Not necessarily. It may mean 80% of eligible text was highlighted. Check the provider's exact definition before interpreting the denominator.

    Is 20% AI bad?

    There is no universal “bad” percentage. Turnitin's 20% rule concerns how it displays lower-confidence results; it is not an academic-misconduct threshold.

    What is the difference between confidence and probability?

    Providers sometimes use the words loosely. A probability should describe a defined event and ideally be calibrated; confidence may be an internal score or label strength. Documentation must explain the usage.

    Can I compare percentages from two detectors?

    Only after confirming they measure the same unit on the same eligible text. Often they do not, which is why averaging scores is usually meaningless.

    Sources and further reading

    Continue the topic

    Related Genutext guides

    View every guide

    Apply the guidance

    Run a first-pass check, then review the context

    Use the free AI preview for a short sample, or sign in for longer AI and plagiarism scans with sentence-level context.