Evaluation Finds No AI Detector Perfectly Accurate
A study of eight AI detection tools identifies Copyleaks as most effective but warns against using scores as definitive verdicts in disciplinary actions.
A comprehensive evaluation of eight AI detection tools found that no single detector is perfectly accurate. Testing involving 480 scans of 60 documents, including academic essays and résumés, revealed that tools struggle most with math, code, and creative writing. While Copyleaks emerged as the most effective tool overall and GPTZero as the best free option, results showed that ChatGPT-generated text remains the hardest to identify, whereas DeepAI is the most easily detected.
Experts warn that these tools frequently produce false positives, particularly when analyzing text from non-native English speakers or human-written content polished by AI. Due to these inaccuracies, Jonathan Gillham, CEO of Originality.ai, advises that detectors should be treated as evidence rather than definitive verdicts. He suggests that AI detection must be part of a broader integrity stack that includes clear usage policies.
Further analysis indicates that detector accuracy depends heavily on training data. If a detector has not encountered a specific AI generator's samples during its training phase, its accuracy can drop significantly, potentially approaching the level of random guessing. Consequently, specialists recommend against using a single detection score for disciplinary action without considering additional context, such as a writer's draft history or prior samples.