Score interpretation

    What an AI-Detection Percentage Means (Is 30% Bad?)

    Distinguish text coverage, probability, confidence and product scores before translating a number into a real-world decision.

    Genutext Editorial Team17 min read
    7 detectors comparedWorked 30% example11 FAQs
    On this page
    1. The short answer
    2. Is there an acceptable AI percentage?
    3. Four different meanings of a percentage
    4. What the percentage means on each detector
    5. Coverage versus probability
    6. Confidence is not certainty
    7. How Turnitin defines its percentage
    8. Worked example: a 30% score on a 2,000-word essay
    9. Is there a 30% rule?
    10. What each percentage usually means, from 5% to 80%
    11. What happens after a high score
    12. Outside the classroom
    13. A five-question reading method
    14. How to record a score without overstating it
    15. Frequently asked questions

    An AI-detection percentage does not have one universal meaning. Depending on the product, it may describe the share of eligible text highlighted, a model's confidence, a calibrated probability, or a score mapped to a label.

    Before interpreting the number, ask: percentage of what?

    The short answer

    No single percentage is "bad". On Turnitin the number is the share of qualifying prose that was highlighted; on GPTZero it is a probability that the document is AI-written; on Originality.ai it is the model's confidence. So "30%" means three different things on three tools. None of the detector vendors or UK universities we reviewed publishes an acceptability line: Turnitin hides scores below 20% because false positives are more common there, and university guidance requires evidence beyond the score before anything happens. The rest of this article explains each of those claims with sources.

    Is your score bad? The 10-second answer, by threshold:

    Your scoreIs it bad?Why
    5%NoBelow the range Turnitin even displays; vendors treat it as noise
    10%NoStill inside Turnitin's suppressed asterisk band
    20%Not by itselfThe first number Turnitin shows; where a report becomes worth reading
    25–30%Depends on the toolQuotes and formulaic prose land here innocently; on confidence tools it still leans human
    40%Read the passagesTwo-fifths of eligible prose flagged, or a score that still leans human
    50–60%Worth a proper lookThe band where vendor labels flip toward "AI" — a signal, not a verdict
    80%+Take it seriouslyTurnitin's own heavy-AI band; still requires evidence beyond the score

    Check your own text against the same signals before anyone else does:

    Try the text in context

    Run a free AI-writing signal check

    Paste 300 characters to 350 words. The sample is analysed for a first-pass signal and then discarded.

    Free AI check

    0 / 350 words free

    Analysed, then discarded — never stored

    Local-currency pricingNo subscriptionNever storedInstant results

    Is there an acceptable AI percentage?

    No. There is no universal "safe" or "acceptable" AI percentage, because the number does not measure wrongdoing — it measures how strongly the text matches patterns the detector associates with AI writing.

    The same 30% can be entirely fine in one context and worth reviewing in another:

    • Is 30% AI bad? On its own, no. Depending on the tool it may mean 30% of eligible prose was highlighted, or a 30% model confidence — neither is a 30% chance anyone broke a rule. A quotation-heavy essay, formulaic methods section or permitted grammar tool can all contribute to a score like this.
    • Is there a threshold institutions use? Some tools suppress or caveat scores below roughly 20% because low readings are unreliable. That is a reliability floor, not an official acceptability line — most universities deliberately publish no threshold, and their guidance says a detector score should not be the sole evidence of misconduct. We cover the UK-specific version of this question in what AI score is acceptable at UK universities?
    • What actually matters? Which sentences were flagged, whether the assignment allowed AI assistance, and what the drafts and version history show. A score only tells you where to look.

    If you are reading a specific number right now, the sections below explain what that percentage is a percentage of — which changes everything about how to read it.

    Four different meanings of a percentage

    1. Text coverage

    The number may be the proportion of eligible prose classified as likely AI-generated. A 30% result then means about 30% of the qualifying text was highlighted—not a 30% chance of misconduct.

    2. Document probability

    A tool may estimate the probability that a document belongs to an AI-generated class. This is a model claim about a class under its calibration assumptions.

    3. Model confidence

    Some systems expose an internal confidence score. Confidence can be high and still wrong, especially when the text differs from the model's validation data.

    4. A normalised product score

    A provider may combine passage outputs into a 0–100 scale and map ranges to “human”, “mixed” or “AI”. Unless the methodology says it is a probability, do not read it as one.

    What the percentage means on each detector

    The fastest way to see why one number cannot be compared with another is to read each vendor's own definition. These are taken from the tools' published pages as of 23 August 2026 (links in the sources list); check the live page before relying on a detail, because vendors change them.

    ToolThe number is a percentage of…What the tool showsMinimum textVendor's own caveat
    Turnitin AI writing reportQualifying long-form prose classed as likely AI-generated or AI-generated-then-paraphrasedHighlights plus an overall %; 1–19% shown as an asterisk300 words of proseShould not be the sole basis for adverse action; independent of the similarity score
    GPTZeroThe probability that the document (and each sentence) is AI-generatedHuman / Mixed / AI label with a confidence categoryNot stated as a word countResults should not be the only proof of AI use
    Originality.aiThe model's confidence: "60% Original" means 60% confident the text is human-written, not a 60/40 split of the textOriginal vs AI confidenceThe score is a confidence, not a proportion
    CopyleaksThe share of the content judged likely to be AI-generated, with separate confidence fields in its APIHighlighted passages and an overall %255 charactersTreat as a signal for review
    QuillBotThe likelihood that the text was generated by AIA single %80 wordsThe tool describes itself as leaning towards "human"
    ScribbrHow much of the text was written or refined by AI, in four categoriesCategory plus %Up to 1,200 words per checkInformational, not proof
    PangramThe proportion of AI-written text, with a separate confidence and a spectrum labelProportion, confidence, labelLabels the middle of the spectrum as mixed or AI-assisted

    Two tools can therefore show "30%" for the same essay and mean, respectively, "30% of the prose is highlighted" and "30% confident it is AI" — which leans human. That is why the reasons detectors disagree matter more than the headline number.

    Coverage versus probability

    Consider a 2,000-word document containing 1,600 words of qualifying prose. If a detector highlights 400 eligible words, it might report 25% because 400 is one quarter of 1,600.

    That number does not mean:

    • there is a 25% chance the student used AI;
    • 25% of the entire file was written by ChatGPT;
    • the writer is 25% responsible for misconduct; or
    • the evidence satisfies a disciplinary threshold.

    The denominator matters because tools may exclude references, quotations, tables, code and short-form material. Compare highlighted passages with the exact report definition.

    Confidence is not certainty

    In machine learning, a confident output can still be a false positive. Calibration tells you how often similar confidence values are correct across a suitable test set; it does not reveal the origin of a particular document.

    Suppose a model is well calibrated and gives many documents a 90% probability. Across comparable documents, roughly nine in ten may belong to the predicted class. The remaining one in ten can still be wrong—and performance may shift for a new language, genre, model or editing process.

    This is why benchmark context matters more than a polished dial. Read how accurate AI detectors are for precision, recall and base-rate examples.

    How Turnitin defines its percentage

    Turnitin's March 2026 guide defines the overall AI percentage as the proportion of qualifying long-form prose its model determines could be AI-generated or AI-generated and then modified by paraphrasing or bypass tools. The full anatomy of the report — who sees it, the institutional on/off switch, and what it cannot show — is in does Turnitin detect AI?

    Turnitin also:

    • requires at least 300 words of prose;
    • supports specified languages and file types;
    • treats the AI percentage as independent of the Similarity score; and
    • suppresses numeric values below 20% because false positives are more frequent in that range.

    The 20% display rule is a product safeguard, not a universal line between human and AI writing. It should never be reused as a misconduct threshold.

    Whether your own university runs this report at all is a separate question — many UK institutions have switched the AI indicator off. The UK University AI-Detection Policy Tracker links each university's own dated statement.

    Worked example: a 30% score on a 2,000-word essay

    Take a 2,000-word essay submitted to a Turnitin-style report that shows 30%. Reading it with the definitions above:

    • Percentage of what? Of the qualifying prose. Turnitin excludes quotations, bullet lists, tables, references and very short or non-English passages. If 1,700 of the 2,000 words qualify, 30% points to roughly 500 words — typically three or four paragraphs — not to 30% of the file.
    • What does the number measure? The share of that prose the model classed as likely AI-generated or AI-modified. It is not a 30% probability that the student used AI, and it is not 30% confidence.
    • Where are the highlighted passages? This is the only part that can be checked. A literature-review section full of formulaic summary sentences, a methods section written to a template, or text polished with a permitted grammar tool are all common, innocent reasons for highlights to cluster.
    • What would change the decision? Drafts and version history, the assignment's stated AI policy, and the student's ability to discuss the highlighted passages. None of these appear in the percentage.

    The same 30% on a 400-word reflection is a different object again: with 400 qualifying words the model is working near its minimum sample and the highlighted 120 words may be a single paragraph, so the reading above applies with even less weight.

    Is there a 30% rule?

    No published one. The phrase "the 30% rule" circulates in forums and on essay-service blogs as if it were a Turnitin or university setting, but:

    • Turnitin's only documented threshold is the 20% display rule. Scores from 1% to 19% are shown as an asterisk because false positives are more frequent in that range. Turnitin does not publish an "acceptable" percentage and says the report must not be the sole basis for adverse action.
    • No UK university we reviewed publishes a percentage line. Most with a public position have switched the indicator off or say a score cannot be sole evidence; the UK University AI-Detection Policy Tracker links each statement.
    • The nearest thing to a published threshold is a triage rule, not an acceptability line. The University of Melbourne's staff guidance tells markers to focus only on work where more than 20% is predicted as AI-generated — and then requires a second piece of evidence before any allegation.

    If someone quotes a 30% rule at you, ask which document it comes from. In our review, none of the pages asserting it cited one.

    What each percentage usually means, from 5% to 80%

    Read each number through the "percentage of what?" lens, and the same score changes meaning by tool. Each threshold below has its own anchor, so you can link a colleague straight to the number in front of you.

    Interactive

    Read your own score

    Enter the number in front of you and what produced it. The reading updates from the same evidence set out in this article. Nothing you enter leaves the page.

    On a probability or confidence tool this still leans human. It is not a 1-in-3 chance of misconduct; it is a weak signal that deserves a look at the flagged passages.

    Keep your drafts and version history, and rewrite any flagged passages in your own voice rather than using a “humaniser” — that protects you far better than chasing a lower number.

    Next: run your own text through the free AI detector to see which passages carry the signal.

    Is 5% AI detection bad?

    No. On Turnitin, anything from 1% to 19% is inside the suppressed range and is displayed as an asterisk rather than a number, because the vendor itself says false positives are more common there. On a probability tool, 5% is a low likelihood for the whole document. A 5% reading is below the range any vendor treats as meaningful.

    Is 10% AI detection bad?

    No — the same suppressed-range logic applies. On Turnitin a 10% result is shown as an asterisk, not a figure; on a probability-based tool it is a low document-level likelihood. It is not evidence of anything on its own, and it does not mean 10% of the essay "is AI".

    Is 20% AI detection bad?

    Not by itself. 20% is the first value Turnitin displays numerically, and the focus threshold in the University of Melbourne's published staff workflow. It marks where a report becomes worth reading, not where anything is proven. Look at which passages are highlighted before drawing any conclusion.

    Is 25% or 30% AI detection bad?

    This is the most-searched band, and the honest answer is: it depends entirely on the tool and the text. On a coverage tool, a quarter to a third of the qualifying prose was highlighted — quotations, formulaic sections and templated writing all land here innocently. On a confidence tool, 25–30% AI still leans human. See the worked 30% example above for how the same number changes meaning.

    Is 40% AI detection bad?

    On a coverage tool, roughly two-fifths of the qualifying prose is highlighted; on a confidence tool, 40% AI still leans human. The highlighted passages decide which reading applies — a literature review full of standard definitions reads very differently from an unbroken run of flagged argument.

    Is 50% or 60% AI detection bad?

    This is the range in which vendor labels tend to switch from "mixed" to "AI": GPTZero's mixed class and Pangram's spectrum labels both sit around here. It is a strong signal to read the passages and check the drafts — not a verdict, and not a number any university publishes as an action threshold.

    Is 80% AI detection bad?

    80% and above is Turnitin's own reporting band for heavily AI-written work. In February 2026 the company reported that roughly 15% of essay submissions between October 2025 and February 2026 scored above 80%, up from about 3% in 2023. A score here still needs the same evidence as any other — drafts, version history and a conversation — before any decision.

    What happens after a high score

    A percentage is the start of a process, not the end of one. Published institutional guidance follows a consistent shape:

    • A second form of evidence is required. Drafts, version history, notes, reference trails and a conversation about the work carry the weight; the score alone does not. Turnitin and GPTZero both say their output should not be the sole basis for a decision.
    • You may not be able to see the score yourself. Turnitin's AI indicator is shown to staff, and Melbourne's guidance states it is not made visible to students. If you have been told a number, ask for the highlighted report and the definition being used.
    • Innocent causes of high scores are well documented. Formulaic sections (methods, literature summaries), quotations and templated text, writing polished with a permitted grammar tool, and short samples all push scores up. The highlighted passages will usually show which of these applies.

    For UK students facing an allegation, our guide to proving you did not use AI covers the evidence that has worked in practice.

    Outside the classroom

    Employers, publishers and content teams set their own rules, and the ones we could find rarely name a percentage at all; where a client imposes a limit it is a contractual choice, not a property of the detector. Google's published guidance on AI-generated content is framed around whether content is helpful and reliable, not around a detector score. The vendor caveats above apply unchanged: a percentage indicates where to look, and the decision remains a human one.

    A five-question reading method

    1. What is the unit? Coverage, probability, confidence or product score?
    2. What text was eligible? Whole file, prose only, sentences or passages?
    3. What is the threshold? Which score becomes a label, highlight or alert?
    4. What validation applies? Same language, genre, length and model version?
    5. What decision is being considered? Screening, conversation or formal action?

    Then read the highlighted text and limitations. If a provider does not explain its score, the number should carry less weight.

    Genutext labels its output as an AI-writing signal and directs users to methodology and limitations. It does not describe the score as proof of authorship. The AI detector page sets out what the free preview and a paid scan each report before you apply the five questions above.

    How to record a score without overstating it

    If a percentage is going into a review note, a marking record or a misconduct file, the wording matters as much as the number. A useful record lets another person understand what was tested and why the result mattered:

    • Record the product, date and any version information shown with the result.
    • Identify the exact text or document section that was submitted.
    • Note whether the input was pasted or extracted from a PDF, and whether the extraction was checked.
    • Write "the detector returned" or "the report classified" — never "the detector proved".
    • Keep the numeric result and its passage-level context together rather than quoting the percentage alone.
    • List the other evidence reviewed, including anything that challenged the concern.
    • Record the policy-based human decision separately from the automated result.

    Frequently asked questions

    Does 80% AI mean an 80% chance the text is AI-written?

    Not necessarily. It may mean 80% of eligible text was highlighted. Check the provider's exact definition before interpreting the denominator.

    Is 20% AI bad?

    There is no universal “bad” percentage. Turnitin's 20% rule concerns how it displays lower-confidence results; it is not an academic-misconduct threshold.

    Is 30% AI bad?

    Not by itself. Depending on the tool, 30% may mean about a third of eligible prose was highlighted or a modest model confidence — neither is a 30% chance of misconduct. Quotation-heavy or formulaic writing can produce a number like this legitimately. What matters is which passages were flagged and whether the assignment allowed AI assistance.

    What percentage of AI detection is acceptable?

    There is no official acceptable percentage. Most universities deliberately publish no threshold, and tools that suppress low scores (such as Turnitin below 20%) do so for reliability, not as an acceptability line. Treat any score as a prompt to review passages and process evidence rather than a pass/fail mark.

    What is the difference between confidence and probability?

    Providers sometimes use the words loosely. A probability should describe a defined event and ideally be calibrated; confidence may be an internal score or label strength. Documentation must explain the usage.

    What is the 30% rule for AI detection?

    An informal shorthand, not a rule. No detector vendor or university document we reviewed defines 30% as a limit. Turnitin's only published threshold is that scores below 20% are shown as an asterisk because false positives are more frequent there.

    Is 25% AI detection bad?

    Not on its own. On Turnitin, 25% means a quarter of the qualifying prose was highlighted — quotations, formulaic sections and permitted grammar tools all push scores into this band. On a confidence tool, 25% leans clearly human. Read the highlighted passages and check what the number is a percentage of before treating it as a concern.

    Is 50% AI detection bad?

    50% sits in the band where tools switch labels from "mixed" toward "AI", so it deserves a proper look at the flagged passages and the drafts. It is still not proof: no vendor describes 50% as a misconduct threshold, and UK guidance requires evidence beyond the score.

    Is 40% AI bad?

    It depends on the unit. On Turnitin it means about 40% of the qualifying prose was highlighted; on Originality.ai it means the model is 40% confident the document is AI-written, which leans towards human. Read the flagged passages, not the number.

    Can students see their Turnitin AI score?

    Usually not. The AI indicator and highlighted report are shown to staff, and the University of Melbourne's guidance states the indicator is not made visible to students. Ask your institution how the report is used and whether you can see the highlights.

    Is the AI percentage the same as the similarity score?

    No. Turnitin states the AI writing percentage is different from, and independent of, the similarity score. A 30% AI score and a 30% similarity score measure unrelated things.

    Does a 10% score mean 10% of my essay is AI?

    Only on coverage-based tools, and Turnitin does not display values below 20% at all. On probability- or confidence-based tools, 10% is a low likelihood for the whole document.

    Can I compare percentages from two detectors?

    Only after confirming they measure the same unit on the same eligible text. Often they do not, which is why averaging scores is usually meaningless.

    Sources and further reading

    Continue the topic

    Related Genutext guides

    View every guide

    Reading a number right now?

    Check your own percentage in context

    Run your text through a scan that shows which sentences carry the signal — not just a headline number. Your first full scan is free when you create an account.