AI Assessment Models Favor AI-Generated Content Over Human Writing
Research reveals generative AI models exhibit systemic bias by awarding higher scores to AI-generated text than to human-written submissions.
Research indicates that generative AI models used to evaluate written submissions, including resumes and research papers, exhibit a systemic bias toward AI-generated content. This preference occurs because models reward stylistic patterns and conformity that mirror their own training data, often resulting in higher scores for AI-written text than for handcrafted human writing.
A July 2026 study titled “Stop Automating Peer Review Without Rigorous Evaluation” identifies a phenomenon called "paper laundering." In these instances, zero-shot large language model rewrites can significantly boost review scores without improving the underlying scientific results. This creates a failure mode where stylistic modifications alone increase the perceived quality of a submission.
This bias creates a paradox for applicants and researchers. Those who adhere to rules prohibiting AI usage may be penalized with lower scores, while those who employ AI to improve their stylistic presentation risk disqualification. The study notes that AI detection tools intended to prevent this practice are unreliable and prone to false positives.