No, they are not reliable enough to be used as evidence, and the false positives are a documented, serious problem. Several major vendors — including OpenAI, which withdrew its own classifier — have effectively conceded this.
Why they fail, mechanically. Detectors mostly measure how *predictable* text is. AI output tends toward statistically likely word choices, so low surprise across a document is treated as a signal. The problem is that plenty of human writing is also highly predictable: formal academic prose, writing by non-native English speakers using simpler constructions, technical writing with fixed terminology, and anyone taught to write in a clear, conventional style. These groups get flagged disproportionately — non-native speakers most of all, which is a well-documented finding and an equity problem, not a rounding error.
Meanwhile, the false negatives are trivially easy: lightly editing AI output, asking for a different style, or running it through a paraphraser defeats most detectors. So the tool is simultaneously punishing some honest students and failing to catch the behaviour it targets — the worst possible combination.
What the scores actually mean: a '80% AI' result is not a probability that it was AI-written, however it's presented in the interface. It's a statistical impression with no calibrated meaning, and no vendor can tell you the false positive rate on your particular student population.
If your friend has been flagged, the practical response:
1. **Produce process evidence.** Version history in Google Docs or Word, drafts, notes, browser history, timestamps. This is far more persuasive than arguing about the detector, and it's why keeping drafts is now genuinely worth doing.
2. **Ask for the institution's policy** on detector evidence, and specifically whether a score alone can support an allegation. Many institutions have quietly moved to 'may not be used as sole evidence' precisely because of the accuracy problem.
3. **Offer to discuss the content.** Someone who wrote a piece can explain their argument, their sources and their choices. That conversation is the fairest test available.
4. **Cite the vendor's own limitations.** Turnitin has publicly acknowledged false positives; OpenAI shut its detector down for low accuracy. These are the vendors' own statements, not partisan claims.
The broader point worth making to institutions: detection is a losing arms race. Assessment designs that are resistant by construction — in-person components, oral defence, drafts and process, work grounded in specific class discussion — solve the problem in a way no classifier will.