As artificial intelligence has grown in many schools, many teachers and administrators are turning to AI detection softwares such as Turnitin to determine whether students wrote their own assignments or the information was fabricated. Programs built into plagiarism checkers promise to identify essays written by artificial intelligence.
In theory, these tools sound like the perfect solution to a growing problem. In reality though, schools are becoming increasingly reliant and a serious concern is emerging; is the technology as accurate as people believe?
“I was scared that it was gonna ruin my reputation and my Cambridge diploma,” said Junior Rajiv Preetam. He is one of many of the juniors who were accused of using artificial intelligence on his global perspectives essay.
For students in high school, the stakes are incredibly high. If an essay is flagged as AI-generated, it can lead to accusations of academic dishonesty resulting in possibly disciplinary meetings. For students in our AICE Global Perspectives course, an accusation of that nature may result in the student’s removal from the Cambridge program, making them ineligible for the AICE Diploma. Yet researchers are increasingly questioning whether detectors are reliable enough to justify those consequences.
Many companies like Turnitin claim over a 98% accuracy rate. However, these numbers often come from controlled tests that do not reflect how students actually write. In the experiments, detectors compare completely AI-generated text with text that has grammatical errors and improper syntax, clearly human-written. Under these obvious conditions, the tools perform well.
Real student writing, however, is far more complicated as students brainstorm ideas, revise drafts, fix grammar, and ask for feedback. When writing with edits is inputted to the detector, it becomes increasingly difficult for the software to determine whether or not a human wrote the text.
Independent studies suggest that the accuracy of these detectors drops significantly when tested under realistic conditions. A study conducted by Mike Perkins et. al.. In 2024 suggested that the accuracy of these detectors drops significantly when tested under realistic conditions. They found that through multiple AI detection tools, they had an average accuracy rate of only 39 percent when evaluating complex writing samples.
In a study conducted by Kalpesh Krishna et al., small edits to writing dropped detection rates from around 70 percent to less than 5 percent. “It’s difficult to balance use versus misuse and how to grade it effectively,” said Social Studies teacher Kelsey Bear.
An even bigger concern is the possibility of false positives, when a human-written essay gets incorrectly labeled as AI-generated. Companies like turnitin claim that these errors happen less than once percent of the time, but some studies suggest this isn’t true. A range of studies suggested that several tools found false positive rates between roughly 5 and 15 percent.
When millions of assignments are processed every year, even a small error rate can affect a large number of students and this is a risk we should not be taking. A detector that incorrectly flags just a few percent of essays could still drastically affectkj hundreds of students who did nothing wrong.
