AI detection and academic integrity

Why detector scores make poor evidence in an academic misconduct case, what a false positive costs a student, and what assessment practices hold up better.

By didaiwritethis.ai

The problem with a percentage

An AI detector returns something like "87% likely AI-generated". That number looks like a measurement. It is closer to a similarity score — an estimate of how much the writing resembles statistically average prose. Those are different claims, and a misconduct process that treats the first as if it were the second will produce false accusations at a predictable rate.

This is not an argument that students never misuse these tools. It is an argument about what a specific piece of evidence can support, and detector output cannot bear the weight institutions routinely put on it.

The arithmetic nobody runs

Suppose a detector has a 1% false positive rate — better than most vendors can demonstrate in independent testing. A department processes 5,000 essays a term, and suppose 2% involve undisclosed AI use.

So roughly a third of everything flagged is innocent — from a tool performing at its advertised best. Lower the true rate, or use a tool with a more realistic 3–5% false positive rate, and flagged-but-innocent work becomes the majority of what lands on your desk.

Nothing about this is exotic. It is the same base-rate problem as any screening test for a relatively uncommon condition, and it is why screening tests are followed by confirmatory tests in every field that takes them seriously.

The failures are not evenly distributed

If false positives fell randomly, this would be a manageable nuisance. They do not. Liang and colleagues (2023) found that GPT detectors misclassified writing by non-native English speakers as AI-generated at dramatically higher rates than writing by native speakers.

The reason is mechanical rather than malicious. Detectors flag low-perplexity text — plain vocabulary, conventional sentence structure. That describes a great deal of competent writing by people working in a second language. It also describes students taught to write in a formal academic register, students on the autism spectrum whose prose runs more structured, and anyone drilled toward a rubric.

An institution relying on these scores is therefore not running a slightly noisy process. It is running one whose errors concentrate on international students and others already less equipped to contest an accusation.

What a false accusation costs

The asymmetry here deserves stating plainly. A missed case of AI use costs a grade that was not earned. A false accusation puts a student through a misconduct process, threatens visa status for international students, and asks them to prove a negative about their own thinking. Those are not comparable errors, and a policy tuned to minimise the first while ignoring the second has its priorities inverted.

Students accused on detector evidence also face a structural problem: there is no clean way to disprove it. "I wrote this" is unfalsifiable from the other side, which is precisely why the burden should not land there.

What holds up better

Assess the process, not just the artifact

Version history is the single most useful thing available, and it is free. Requiring submission through a platform that retains revision history — Google Docs, Word with AutoSave, a repository — turns an unanswerable question into an answerable one. A document that grew over two weeks looks nothing like one pasted in whole.

Build in checkpoints

Proposal, annotated bibliography, draft, revision. Each stage is low-stakes on its own, makes wholesale outsourcing impractical, and produces exactly the evidence trail that settles disputes later. It also happens to be better pedagogy.

Talk to the student

A short conversation about the work — why this argument, what got cut, what nearly went in instead — is more diagnostic than any score, and it is fair in a way a percentage is not, because it gives the person a chance to respond. Where a detector flag prompts a conversation rather than a charge, it is being used sensibly.

Say what is actually allowed

A great deal of suspected misconduct is genuine confusion. "No AI" means different things to different students, and it is rarely clear whether it covers grammar checking, brainstorming, or a tool built into the word processor. Policies that name specific permitted and prohibited uses, per assignment, prevent more problems than any detector catches.

Where watermark detection fits

Watermarking is a genuinely better class of evidence than style-based detection: a positive result indicates a specific model's signal rather than a resemblance to average prose. It does not resolve the problem for academic use, for three reasons:

Useful as one input among several. Not a replacement for knowing how a piece of work came together. See how watermarking works for the mechanism, and the evidence on style-based detectors for the comparison.

A reasonable position

Use detector output, if at all, as a prompt to look more closely — never as the finding itself. Require process artifacts so that looking more closely is possible. Write policies that say what is permitted rather than gesturing at a prohibition. And when a case does turn on a score alone, recognise that the score is not evidence of what it appears to be evidence of.


Check a piece of text. didaiwritethis.ai looks for a Claude watermark in text you paste — a different method from the style-based detectors discussed here. Anthropic's detection API has not been released yet, so results today are a labelled preview.

try it