A detector flagged a student. Now what?
Updated 2026-09-02 · 8 min read
You are looking at a percentage and deciding what to do about a person. The number is not a record of what the student did. It is an estimate of how closely their sentences resemble machine-generated prose, produced by a model that never watched anyone write.
What the percentage actually counts
Turnitin splits the document into sentences, scores each one, and reports the share that crossed an internal threshold. Quotations, reference lists and very short sentences are excluded, so the denominator is smaller than the document in front of you.
A heavily quoted paper can therefore produce a startling figure from a handful of scored sentences. Below 300 words of continuous prose there is no AI report at all: Turnitin raised that minimum from 150 words on the stated grounds that accuracy improves with length. The report also cannot name a tool. It does not separate ChatGPT from Gemini from an accepted Grammarly rewrite from a student who simply writes evenly.
The error rate Turnitin publishes about itself
Annie Chechitelli, Turnitin's chief product officer, said in 2023 that the sentence-level false positive rate is around 4%, and that documents scoring under 20% AI writing show a higher rate of false positives. That is why those scores display with an asterisk. The document-level figure claimed at launch was under 1%.
Run the 1% against your own volume. Vanderbilt did the arithmetic in public in August 2023: roughly 75,000 papers went through Turnitin there during 2022, so a 1% document-level false positive rate implies about 750 papers a year carrying a flag they did not earn. Vanderbilt disabled the detector, citing that arithmetic, the reported over-flagging of non-native English writers, and the fact that Turnitin does not explain how a determination is reached.
The sentence figure deserves more attention than it gets, because highlighted sentences are what most of us actually look at. At 4%, a forty-sentence essay typed by hand will average one or two highlights by chance alone.
Your policy probably says the score is not enough
Turnitin's own guidance to educators says the indicator should not be the sole basis of an academic misconduct finding, and most institutional policies say something to the same effect. Find the sentence in yours before you act.
If it is there, a case resting on the score alone is procedurally defective. It will be overturned on appeal, but only after the student has spent a term frightened, and the overturning is not a happy ending for anyone involved.
Check these before you conclude anything
- Was the whole document scored, or an extract? A figure on a fragment is worth much less.
- How many sentences were actually scored once quotations and references dropped out?
- Is the score under 20%, where Turnitin itself signals reduced reliability?
- Is this a second-language writer? Liang and colleagues, publishing in Patterns in 2023, ran 91 human-written TOEFL essays through seven detectors: most were misclassified as machine-generated, and 19 were flagged by all seven at once.
- Does the writing match what you have seen from this student in class, in earlier drafts, or in anything written under your eye?
How to raise it without opening a case
Ask, do not tell. 'I would like to talk about how you wrote this' starts a conversation a student who wrote the paper can win in five minutes. 'This was flagged as AI' starts something else, and once a referral leaves your hands nobody who picks it up will know the student or the work.
Then ask them to walk you through it: why that source, what they cut, which paragraph fought back hardest. Someone who generated the document cannot do this. Someone who wrote it usually cannot be stopped. A short oral check is the closest thing to direct evidence available here, which is why a growing number of departments have made it a formal step rather than an improvisation.
- Ask for the process record: version history, outline, notes, the sources they read and rejected.
- Ask what tools they used and for what, in specific terms. 'I used Grammarly' covers both a comma fix and a wholesale rewrite.
- Put the meeting before the referral, not after it.
Design that removes the question next term
- Ask for a version history link with every submission, as routine practice rather than as suspicion of anyone.
- Set work that leans on class discussion, local material, or the student's own earlier writing.
- Add a short in-person or recorded component to major pieces.
- Write the tool policy in specific terms. Spelling correction and generated text are different things, and students genuinely cannot guess where you draw the line between them.
The two errors are not the same size
A missed case of AI use costs you one inflated grade in one module. A wrong accusation costs a student weeks of anxiety, sometimes a transcript notation, and for an international student it can reach their visa. It also spends the thing the rest of your teaching runs on, which is that students believe you will be fair to them.
None of this is an argument for ignoring AI use. It is an argument for treating the score as the start of an inquiry you conduct yourself, using evidence a machine cannot produce.
The short version
Read the score as a reason to look, not as a finding. Check the length, the excluded sentences and the 20% asterisk, then ask the student to talk you through the work. Turnitin's guidance says the number should not be the sole basis of a misconduct decision, and your own policy probably says it too.