Score interpretation
What Does an AI-Detection Percentage Actually Mean?
Distinguish text coverage, probability, confidence and product scores before translating a number into a real-world decision.
On this page
An AI-detection percentage does not have one universal meaning. Depending on the product, it may describe the share of eligible text highlighted, a model's confidence, a calibrated probability, or a score mapped to a label.
Before interpreting the number, ask: percentage of what?
Four different meanings of a percentage
1. Text coverage
The number may be the proportion of eligible prose classified as likely AI-generated. A 30% result then means about 30% of the qualifying text was highlighted—not a 30% chance of misconduct.
2. Document probability
A tool may estimate the probability that a document belongs to an AI-generated class. This is a model claim about a class under its calibration assumptions.
3. Model confidence
Some systems expose an internal confidence score. Confidence can be high and still wrong, especially when the text differs from the model's validation data.
4. A normalised product score
A provider may combine passage outputs into a 0–100 scale and map ranges to “human”, “mixed” or “AI”. Unless the methodology says it is a probability, do not read it as one.
Try the text in context
Run a free AI-writing signal check
Paste 300 characters to 350 words. The sample is analysed for a first-pass signal and then discarded.
Coverage versus probability
Consider a 2,000-word document containing 1,600 words of qualifying prose. If a detector highlights 400 eligible words, it might report 25% because 400 is one quarter of 1,600.
That number does not mean:
- there is a 25% chance the student used AI;
- 25% of the entire file was written by ChatGPT;
- the writer is 25% responsible for misconduct; or
- the evidence satisfies a disciplinary threshold.
The denominator matters because tools may exclude references, quotations, tables, code and short-form material. Compare highlighted passages with the exact report definition.
Confidence is not certainty
In machine learning, a confident output can still be a false positive. Calibration tells you how often similar confidence values are correct across a suitable test set; it does not reveal the origin of a particular document.
Suppose a model is well calibrated and gives many documents a 90% probability. Across comparable documents, roughly nine in ten may belong to the predicted class. The remaining one in ten can still be wrong—and performance may shift for a new language, genre, model or editing process.
This is why benchmark context matters more than a polished dial. Read how accurate AI detectors are for precision, recall and base-rate examples.
How Turnitin defines its percentage
Turnitin's March 2026 guide defines the overall AI percentage as the proportion of qualifying long-form prose its model determines could be AI-generated or AI-generated and then modified by paraphrasing or bypass tools.
Turnitin also:
- requires at least 300 words of prose;
- supports specified languages and file types;
- treats the AI percentage as independent of the Similarity score; and
- suppresses numeric values below 20% because false positives are more frequent in that range.
The 20% display rule is a product safeguard, not a universal line between human and AI writing. It should never be reused as a misconduct threshold.
A five-question reading method
- What is the unit? Coverage, probability, confidence or product score?
- What text was eligible? Whole file, prose only, sentences or passages?
- What is the threshold? Which score becomes a label, highlight or alert?
- What validation applies? Same language, genre, length and model version?
- What decision is being considered? Screening, conversation or formal action?
Then read the highlighted text and limitations. If a provider does not explain its score, the number should carry less weight.
Genutext labels its output as an AI-writing signal and directs users to methodology and limitations. It does not describe the score as proof of authorship.
Frequently asked questions
Does 80% AI mean an 80% chance the text is AI-written?
Not necessarily. It may mean 80% of eligible text was highlighted. Check the provider's exact definition before interpreting the denominator.
Is 20% AI bad?
There is no universal “bad” percentage. Turnitin's 20% rule concerns how it displays lower-confidence results; it is not an academic-misconduct threshold.
What is the difference between confidence and probability?
Providers sometimes use the words loosely. A probability should describe a defined event and ideally be calibrated; confidence may be an internal score or label strength. Documentation must explain the usage.
Can I compare percentages from two detectors?
Only after confirming they measure the same unit on the same eligible text. Often they do not, which is why averaging scores is usually meaningless.