Notes

AI did not break assessment. It exposed it.

The Financial Times reported last week that universities are abandoning AI-detection tools — too many false positives, and the results can’t reliably support misconduct allegations. A recent HEPI piece makes it sharper: the systems may be especially unreliable for international students and people writing in a second language. The tool designed to catch cheating is most likely to flag the students with least institutional protection.

The interesting argument, though, is not simply that the technology is unreliable. It is that universities tried to use software to answer an educational question.

Did this student actually understand and produce this work?

A probability score cannot answer that. It can only generate suspicion — and suspicion distributed unevenly, along lines that map uncomfortably onto existing disadvantage.

The deeper problem is that detection was always a patch on a design flaw. If an assessment can be completed adequately by a language model, the assessment was not measuring what it claimed to measure. It was measuring the ability to produce a certain kind of text. AI did not create that problem. It made the problem visible and cheap to exploit at scale.

Universities that are now asking how to detect AI use are asking the wrong question first. The right question is what they actually want to know about a student’s learning — and whether their current assessments can tell them that. Some can. Many cannot, and they were not very good at it before ChatGPT either.

Detection is the wrong frame. Redesign is the harder one, which is probably why it keeps getting deferred.

Via FT: Universities abandon AI detectors