The short answer
- A detector score is not evidence. Detectors are unreliable and biased against non-native English writers — treat a flag as a prompt to look, never proof.
- The most reliable signal is knowing your students — voice, in-class writing, and the homework-vs-test gap beat any software.
- Never open with an accusation. Ask the student to explain their work; understanding can’t be faked, and the conversation is fairer than a score.
- Document your reasoning if it escalates — process artifacts and your own observations, not a detector percentage.
- The best “detection” is prevention through assignment design.

Sooner or later a submission lands on your desk that doesn’t sit right, and the temptation is to reach for a detector and let it decide. Please don’t let it. Detectors are wrong often enough — and wrong in a biased way — that a score can turn a hunch into an injustice. Here’s how to handle suspicion fairly, in a way that survives a parent meeting and doesn’t burn a student who didn’t do anything.
Why isn’t a detector score evidence?
Because it estimates predictability, not authorship — and it’s wrong too often to trust. An AI detector measures how statistically ordinary a text is; plenty of honest human writing is ordinary. The full evidence is here, and the headline is stark: in a Stanford study, seven detectors falsely flagged human-written essays by non-native English speakers 61% of the time on average, while flagging native-born students at near zero. Even a vendor’s best-case 1% false-positive rate produces roughly 15 wrong accusations across a term of 1,500 submissions. And detector companies themselves say a score isn’t proof.
So a flag is a reason to look, never a reason to conclude. Build your judgment on better signals.
What signals are actually reliable?
The ones that come from knowing your students. No software has these:
- Voice. You know how each student writes and talks. Work that’s suddenly far more fluent, formal, or generic than their established voice is worth a second look — and unlike a detector, your knowledge of a specific student isn’t biased against a whole demographic.
- The homework–test gap. Spotless homework, sinking test scores. This is the most durable tell there is, because it reflects the actual problem: the homework was outsourced, the test wasn’t.
- Specifics that aren’t there. Work with no reference to your class, no personal example, no trace of the process — generic in a way a student who was there usually isn’t.
- Unverifiable content. Invented citations and facts that don’t check out are a strong signal of unedited AI.
How should you handle it?
Open with a question, not a charge. The single most reliable and fairest move is to ask the student to explain their work — the argument, the choices, where a fact came from, what they’d change. Understanding can’t be faked on the spot: a student who did the thinking walks you through it; one who didn’t stalls. This conversation is more accurate than any score, and it treats the student as someone you’re trying to understand rather than prosecute.
| Symptom | Likely cause | Response |
|---|---|---|
| High detector score, writes this way in class | False positive | Disregard the score |
| High score, non-native English writer | Documented detector bias | Do not act on the score |
| Can’t explain their own argument | Possible outsourcing | Teach the honest approach; look at process |
| Flawless homework, low test scores | Outsourced practice | Address the learning gap directly |
| Everything’s fine but it “feels off” | Trust it enough to ask, not to accuse | Have the conversation |
The better answer is upstream
The most effective “detection” is not needing to detect. Assignments designed to be AI-resistant — process artifacts, personal connection, in-class components — surface who did the work as a byproduct, without anyone playing forensic analyst. Pair that with a clear AI policy where honesty is the norm, and the number of submissions that “don’t sit right” drops on its own.
Detection will always be part of teaching. Just make it the human kind — knowing your students, asking good questions — rather than outsourcing your judgment to software that’s biased and wrong.
Frequently asked questions
Can teachers reliably detect AI-written homework?
Not through software. AI detectors produce false positives and negatives and can’t prove authorship, so they’re unreliable as evidence. Teachers detect AI use far more reliably through knowledge of a student’s usual voice, the gap between homework and test performance, and — most of all — asking the student to explain their work, which a genuine author can do and an outsourcer can’t.
Can I fail a student based on an AI detector score?
You shouldn’t rely on a score alone, and many institutions require corroborating evidence for integrity cases. Detector vendors themselves caution that scores aren’t proof. A defensible case rests on process artifacts, a student’s inability to explain their own work, and documented observations — not a probability from a tool that’s wrong too often to trust.
What should I do if I suspect a student used AI?
Start with a conversation, not an accusation. Ask them to walk you through their reasoning, their sources, and their choices. A student who did the work can; one who outsourced it usually can’t. Approach it as trying to understand what happened and to teach, rather than to prosecute — it’s fairer and, with false-positive risk so high, safer.
Why are AI detectors biased against some students?
Most detectors flag text for being statistically predictable, and writing in a second language tends to use common vocabulary and simpler structures that read as predictable. A Stanford study found detectors falsely flagged non-native English writers’ essays 61% of the time on average, while barely flagging native writers — a bias that’s a direct result of how the tools work.