Score interpretation

    Why Do AI Detectors Give Different Scores for the Same Text?

    Compare models without averaging unlike percentages, using a practical checklist for thresholds, eligible text, dates and score definitions.

    Genutext Editorial Team5 min read
    7 causes4 score formats8-step checklist
    On this page
    1. Seven reasons scores differ
    2. The percentages may mean different things
    3. Why the same tool can change its result
    4. Should you average detector scores
    5. A practical comparison checklist
    6. Frequently asked questions

    AI detectors disagree because they are not measuring with one shared model, dataset, threshold or score definition. Two tools can analyse identical text and answer slightly different statistical questions.

    A result of 70% in one product may mean “70% of eligible prose was highlighted”. In another, it may mean “the document falls in a high-likelihood category”. Those numbers look comparable but may not share a denominator.

    Seven reasons scores differ

    1. Different training data

    Each detector learns from its own collection of human and AI-generated examples. Academic essays, marketing copy, technical prose, second-language writing and newer language models may be represented differently. A detector usually performs best on material similar to its validation data.

    2. Different model architecture

    Some systems classify whole documents; others combine sentence or passage classifications. Providers may use ensembles, linguistic features, language-model probabilities or proprietary representations. The final outputs need not react to the same patterns.

    3. Different thresholds

    A cautious detector may require strong evidence before flagging text, reducing false positives but missing more AI writing. A more sensitive threshold may find more AI passages while also increasing false alarms. That is a design trade-off, not simply a “better” or “worse” score.

    4. Different eligible text

    Tools may exclude references, quotations, tables, code, lists or very short passages. Turnitin, for example, defines qualifying text as long-form prose and requires at least 300 words for its AI report. Another checker may accept a much shorter input.

    5. Different preprocessing

    File conversion, line breaks, headers, OCR errors, punctuation and language detection can change the text reaching the classifier. Copying from a PDF may not be equivalent to uploading the original document.

    6. Different model versions

    Detectors are updated as language models evolve. Turnitin's release notes explicitly say old reports are not recalculated and resubmission is required to use a new model. A comparison must record the date and version where available.

    7. Normal uncertainty

    Text near a decision threshold can move category after a small edit. Mixed human/AI work, formulaic language and unfamiliar genres often produce less stable evidence than long, unedited samples that resemble training data.

    Try the text in context

    Run a free AI-writing signal check

    Paste 300 characters to 350 words. The sample is analysed for a first-pass signal and then discarded.

    The checker loads as you approach it.

    The percentages may mean different things

    Before comparing outputs, identify the unit:

    Output wordingPossible meaningQuestion to ask
    “35% AI”Share of eligible text highlightedWhich parts of the document were eligible?
    “80% probability”Model confidence or calibrated probabilityProbability of what event, at what level?
    “Likely AI”Category after a thresholdWhat threshold and error trade-off apply?
    Sentence highlightsLocal passage classificationsHow are they combined into the document result?

    Do not translate one format into another without provider documentation. In particular, “80% AI” does not automatically mean an 80% chance that the writer used AI.

    For a fuller guide, read what an AI-detection percentage actually means.

    Why the same tool can change its result

    Even within one product, results can change because:

    • the provider released a new model;
    • the submitted text or file conversion changed;
    • more or less surrounding context was included;
    • the supported-language model changed;
    • a borderline result moved across a threshold; or
    • the service corrected a processing bug.

    Keep the original input, date and report when a result informs a serious decision. A screenshot of a percentage without the eligible text or model context is weak evidence.

    Should you average detector scores

    Usually not. Averaging 20%, 60% and 90% produces 57%, but that number has no defined meaning if the tools use different units and thresholds.

    Running more detectors can also create confirmation bias: the reviewer may keep scanning until one tool supports the suspicion. Multiple outputs are useful only when the comparison method is set in advance and the reports are interpreted alongside independent evidence.

    For academic review, one documented detector result plus drafts, citations, version history and a student conversation is generally more informative than a row of unexplained percentages.

    A practical comparison checklist

    1. Use exactly the same plain text where possible.
    2. Confirm each tool supports the language, genre and length.
    3. Record the date, model version and settings.
    4. Define what each percentage or label means.
    5. Compare highlighted passages, not only headline scores.
    6. Note exclusions such as references or quotations.
    7. Treat disagreement as uncertainty, not as evidence of guilt.
    8. Return to the relevant policy and writing-process evidence.

    Genutext presents its free result as an AI-writing signal and links to its methodology and limitations so the score is not mistaken for a disciplinary conclusion.

    Frequently asked questions

    Which AI detector should I trust when scores disagree?

    Start with the tool's documented scope, validation and score definition. If the decision affects a person, do not choose the most severe result; review independent evidence and allow an explanation.

    Does disagreement prove AI detectors are useless?

    No. It shows that outputs are model-dependent and uncertain. A detector can support screening or review without being reliable enough to act as proof.

    Why did Turnitin and another detector give opposite answers?

    They may differ in training data, eligible text, thresholds, supported languages and output definitions. Turnitin's percentage specifically relates to qualifying prose in its report.

    Can copying from a PDF change the score?

    Yes. OCR errors, lost punctuation, headers and line breaks can alter the classifier input. Compare the actual extracted text before treating two runs as equivalent.

    Sources and further reading

    Continue the topic

    Related Genutext guides

    View every guide

    Apply the guidance

    Run a first-pass check, then review the context

    Use the free AI preview for a short sample, or sign in for longer AI and plagiarism scans with sentence-level context.