Two detectors, one paragraph, two verdicts
Updated 2026-08-28 · 7 min read
The disagreement is not a malfunction. Each tool learned from a different corpus, draws its cutoff somewhere else, and reports a different quantity, so there was never a reason for two numbers to match. Which of those is driving your particular split decides what you can do with it.
The percentages are not the same quantity
Before you compare two numbers, check that they measure the same thing. Often they do not.
Turnitin reports the share of qualifying sentences that crossed an internal threshold. Some tools report a probability that the document as a whole was machine-generated. Others report a confidence in their own verdict, which is a third thing again, and a few report a score with no published definition at all.
So 40 from one tool and 40 from another can mean 'four sentences in ten crossed a line' and 'the model is moderately sure about the whole document'. Those are not comparable. Averaging them, or treating the higher one as the real answer, produces a number that means nothing.
Four reasons two honest tools split
- Training data. A detector learns what machine text looks like from whatever it was shown. One built mainly on older GPT output will score text from a newer model, or from Claude or Gemini, differently.
- Threshold. Every tool picks a cutoff for calling text AI, and where it sits is a commercial decision as much as a technical one. Move the line and the same document changes verdict while the model stays identical.
- Segmentation. Scoring sentence by sentence and scoring a whole document give different answers on the same page, especially when one or two paragraphs are unusual.
- Short-text handling. Below a few hundred words every one of them becomes unreliable, and they become unreliable in different directions.
Length is the biggest single cause of a wild split
If you are comparing scores on one paragraph rather than a full paper, expect chaos. These systems estimate a statistical property of a text, and a statistical property of ninety words is barely estimable at all.
Try it: add one sentence and run the same passage again. The number moves. That instability is not evidence of anything sinister, it is what estimating from a small sample looks like.
It has a practical consequence worth raising. A score on an extract is worth much less than a score on the whole document, so if only part of your work was tested, ask for the figure on the full text.
The demonstration that works, and the one that backfires
The tempting move is to run your flagged paper through two more detectors, find the one that says zero, and screenshot it. It is weak, and it can hurt you. The obvious reply is that one of your three tools still says you cheated, and you have volunteered two new numbers into a conversation you wanted to be about your drafts.
The stronger version tests the instrument instead of your paper. Take writing that is unarguably human and unarguably not yours to fake: something you wrote years before any of these tools existed, a passage from a book on the reading list, or public-domain text. Run it through the same detector, the same day, at the same settings. If it flags, you have shown that this tool produces false positives on known-human writing, using material nobody can accuse you over.
What the disagreement does and does not prove
It proves the tools are estimating rather than observing. If they were reading provenance, meaning a signed record of how a file was actually produced, they would agree, because there would be nothing left to disagree about. C2PA Content Credentials and Google's SynthID work that way for images: two readers of the same manifest return the same answer.
It does not prove that any particular one of them is wrong about you. A detector that flags your paper and a detector that clears it are both guessing, and the second guess is not an acquittal. Be clear-eyed about that before you lean on it, because arguing that disagreement proves your innocence invites the reply that it proves nothing in either direction, and that reply is correct.
How to use it in the meeting
- Mention it once, in a sentence: the tools disagree, which tells you they estimate rather than observe.
- Ask which tool produced the score, what it reported, and whether the institution has a written threshold for acting on it.
- Ask whether the whole document was scanned or only an extract.
- Then go back to your process: version history, outline, the sources you rejected, and an offer to talk through the argument paragraph by paragraph.
The short version
Detectors disagree because they trained on different text, cut at different thresholds, and report different quantities. Say that once, demonstrate it on writing nobody can dispute rather than on your own paper, and spend the rest of the meeting on how you wrote it.