Score interpretation
Why Do AI Detectors Give Different Scores for the Same Text?
Compare models without averaging unlike percentages, using a practical checklist for thresholds, eligible text, dates and score definitions.
On this page
AI detectors disagree because they are not measuring with one shared model, dataset, threshold or score definition. Two tools can analyse identical text and answer slightly different statistical questions.
A result of 70% in one product may mean “70% of eligible prose was highlighted”. In another, it may mean “the document falls in a high-likelihood category”. Those numbers look comparable but may not share a denominator.
Seven reasons scores differ
1. Different training data
Each detector learns from its own collection of human and AI-generated examples. Academic essays, marketing copy, technical prose, second-language writing and newer language models may be represented differently. A detector usually performs best on material similar to its validation data.
2. Different model architecture
Some systems classify whole documents; others combine sentence or passage classifications. Providers may use ensembles, linguistic features, language-model probabilities or proprietary representations. The final outputs need not react to the same patterns.
3. Different thresholds
A cautious detector may require strong evidence before flagging text, reducing false positives but missing more AI writing. A more sensitive threshold may find more AI passages while also increasing false alarms. That is a design trade-off, not simply a “better” or “worse” score.
4. Different eligible text
Tools may exclude references, quotations, tables, code, lists or very short passages. Turnitin, for example, defines qualifying text as long-form prose and requires at least 300 words for its AI report. Another checker may accept a much shorter input.
5. Different preprocessing
File conversion, line breaks, headers, OCR errors, punctuation and language detection can change the text reaching the classifier. Copying from a PDF may not be equivalent to uploading the original document.
6. Different model versions
Detectors are updated as language models evolve. Turnitin's release notes explicitly say old reports are not recalculated and resubmission is required to use a new model. A comparison must record the date and version where available.
7. Normal uncertainty
Text near a decision threshold can move category after a small edit. Mixed human/AI work, formulaic language and unfamiliar genres often produce less stable evidence than long, unedited samples that resemble training data.
Try the text in context
Run a free AI-writing signal check
Paste 300 characters to 350 words. The sample is analysed for a first-pass signal and then discarded.
The percentages may mean different things
Before comparing outputs, identify the unit:
| Output wording | Possible meaning | Question to ask |
|---|---|---|
| “35% AI” | Share of eligible text highlighted | Which parts of the document were eligible? |
| “80% probability” | Model confidence or calibrated probability | Probability of what event, at what level? |
| “Likely AI” | Category after a threshold | What threshold and error trade-off apply? |
| Sentence highlights | Local passage classifications | How are they combined into the document result? |
Do not translate one format into another without provider documentation. In particular, “80% AI” does not automatically mean an 80% chance that the writer used AI.
For a fuller guide, read what an AI-detection percentage actually means.
Why the same tool can change its result
Even within one product, results can change because:
- the provider released a new model;
- the submitted text or file conversion changed;
- more or less surrounding context was included;
- the supported-language model changed;
- a borderline result moved across a threshold; or
- the service corrected a processing bug.
Keep the original input, date and report when a result informs a serious decision. A screenshot of a percentage without the eligible text or model context is weak evidence.
Should you average detector scores
Usually not. Averaging 20%, 60% and 90% produces 57%, but that number has no defined meaning if the tools use different units and thresholds.
Running more detectors can also create confirmation bias: the reviewer may keep scanning until one tool supports the suspicion. Multiple outputs are useful only when the comparison method is set in advance and the reports are interpreted alongside independent evidence.
For academic review, one documented detector result plus drafts, citations, version history and a student conversation is generally more informative than a row of unexplained percentages.
A practical comparison checklist
- Use exactly the same plain text where possible.
- Confirm each tool supports the language, genre and length.
- Record the date, model version and settings.
- Define what each percentage or label means.
- Compare highlighted passages, not only headline scores.
- Note exclusions such as references or quotations.
- Treat disagreement as uncertainty, not as evidence of guilt.
- Return to the relevant policy and writing-process evidence.
Genutext presents its free result as an AI-writing signal and links to its methodology and limitations so the score is not mistaken for a disciplinary conclusion.
Frequently asked questions
Which AI detector should I trust when scores disagree?
Start with the tool's documented scope, validation and score definition. If the decision affects a person, do not choose the most severe result; review independent evidence and allow an explanation.
Does disagreement prove AI detectors are useless?
No. It shows that outputs are model-dependent and uncertain. A detector can support screening or review without being reliable enough to act as proof.
Why did Turnitin and another detector give opposite answers?
They may differ in training data, eligible text, thresholds, supported languages and output definitions. Turnitin's percentage specifically relates to qualifying prose in its report.
Can copying from a PDF change the score?
Yes. OCR errors, lost punctuation, headers and line breaks can alter the classifier input. Compare the actual extracted text before treating two runs as equivalent.