The teacher says an AI wrote your essay. It didn't. What to do now
An AI detector is not proof: OpenAI pulled its own in 2023 for low accuracy, and a Stanford study watched seven detectors flag more than half of essays written by real people as machine work. What is proof is the version history of your document, and you save it today.
It happens more and more: you hand in an essay you wrote yourself, at your own desk, with your own mistakes, and it comes back with a note in the margin: the detector says it is AI. The first thing to know is that the detector does not know that. None of them do. What it does is calculate a probability from how predictable your text sounds, and a correct, clear text with simple vocabulary sounds very predictable. The better you follow the rules from English class, the more you look like a machine to the software.
What detectors can and cannot do
- OpenAI, the company behind ChatGPT, released its own text detector on 31 January 2023 and withdrew it on 20 July of the same year. By its own tests it recognised only 26 percent of AI-written text, and it labelled 9 percent of human-written text as AI.
- A Stanford University study published in the journal Patterns in July 2023 ran 91 TOEFL essays, written by people whose first language is not English, through seven widely used detectors. On average 61.3 percent were flagged as AI. At least one detector flagged 97.8 percent. Essays by US eighth graders, by contrast, were classified almost perfectly.
- The same study had those essays rewritten with richer vocabulary and the false positive rate dropped to 11.6 percent. In other words, the detector was not measuring whether a machine had written; it was measuring whether the vocabulary was simple.
- Turnitin, the detector most schools and universities use, says of itself that it wrongly flags fewer than 1 percent of documents and around 4 percent of sentences. A school marking a thousand essays a term can therefore, by the maker's own figures, have several wrongly accused students every term.
As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy.
The first 24 hours: save the evidence
- Do not touch the document. In Google Docs open File, Version history, See version history: every session is there with date and time, and you can watch the text grow paragraph by paragraph. In Word, version history only exists if the file was saved to OneDrive or SharePoint: click the file name in the title bar and choose Version History. Take screenshots of the version list.
- Gather everything that existed before the text: notes, the handwritten outline, photos of your notebook, the mind map, the pages you read with your own highlighting.
- Save your browser history from those days: your searches and the pages you opened tell the story of how you researched.
- Ask for the report. Which tool was used, what number it gave, whether it flagged individual sentences or the whole document. A percentage without context is not an accusation, it is a question.
- Offer to talk about the text. Five minutes explaining out loud why you chose that argument and what you found hardest is the best evidence there is, and no software can fake it.
- Do not rewrite anything and do not run your text through another detector to prove the opposite. Two detectors that disagree only prove that detectors disagree.
The meeting with school: five sentences for your parents
- Thank you for making the time. We want to understand what happened, not to argue.
- Which tool flagged the text, and what exactly did it show: a percentage, specific sentences or the whole document?
- We have the version history of the document from the first line to the last and would like to show it to you.
- The maker of the tool itself acknowledges false positives, so we ask that the history counts as evidence in the same way the report does.
- How do we do this from now on so it does not happen again: always writing in a document with a history, a short oral conversation about each piece of work, or both?
And if this time it really was the AI
Then the best move is the most uncomfortable one: say it yourself, early, to the teacher, before the meeting with your parents and before anyone proves it. Ask to redo the work and ask what use of AI is allowed in that subject, because it is almost never written down and can almost always be asked. A student who owns the mistake and hands in a second version of their own leaves that conversation with more credit than they walked in with. A student who defends a lie while the version history shows the whole text appearing in a single paste does not.
The habit that protects you for good costs nothing: always write, from the first line, in a document with version history, and write inside it instead of pasting blocks from elsewhere. Notes, outline and the three attempts at a first paragraph can go at the end of the same file. That way every piece of work arrives with its own biography, and on the day a detector gets you wrong, the proof is already made.
Sources
- OpenAI: New AI classifier for indicating AI-written text (with the withdrawal notice of 20 July 2023)
- Liang, Yuksekgonul, Mao, Wu, Zou: GPT detectors are biased against non-native English writers, Patterns, 10 July 2023
- Turnitin: Understanding the false positive rate for sentences of our AI writing detection capability, 14 June 2023
- Google Docs Help: see the version history of a file
- Microsoft Support: View previous versions of Office files
- Can my child use AI for homework? There is no rule, there is the school's
- AI as a study partner
FAQ
Can an AI detector prove a text was written by a machine?
No. It calculates a probability from how predictable the text sounds. OpenAI withdrew its own detector in July 2023 for low accuracy: by its own tests it recognised 26 percent of AI text and wrongly flagged 9 percent of human text. Turnitin acknowledges false positives in fewer than 1 percent of documents and around 4 percent of sentences.
Why do essays with simple vocabulary get flagged more often?
Because detectors measure how predictable a text is, and a correct text made of common words is very predictable. The Stanford study in Patterns saw seven detectors flag on average 61.3 percent of TOEFL essays by non-native writers; after rewriting with richer vocabulary the rate fell to 11.6 percent.
What is the best evidence that I wrote the text myself?
The document's version history, which shows how the text grew session by session, together with the notes and outline from before and a five minute oral conversation about your own work. Save it the same day and do not touch the document.
What if I really did use the AI?
Tell the teacher early and on your own initiative, ask to redo the work and ask what use is allowed in that subject. Owning it and handing in a second version of your own earns more credit than defending a lie the version history takes apart.
Did this article help you?
The three books in the series
One approach, create instead of consume, in three versions depending on who is reading.
The whole series on WhileYouCreate.com
The site has the free companion pages for all three books, the downloadable material and the journal.
Go to WhileYouCreate.com →Learn AI as a family →

