Detection Accuracy of AI-Generated Academic Text Across Major AI-Detection Platforms: A Comparative Evaluation of Human, Machine-Written and AI-Assisted Manuscripts

Authors

DOI:

https://doi.org/10.66687/JMRIS

Keywords:

authorship attribution, academic writing, false positives, Chat GPT

Abstract

Background: Artificial-intelligence text detectors are increasingly used in universities and scholarly publishing to estimate whether prose was generated by a large language model. Their outputs can influence academic-integrity investigations despite uncertainty regarding false positives, model drift, language bias and robustness to editing.
Objective: To compare the performance characteristics of major AI-text detection platforms in academic writing and identify conditions under which detector scores become unreliable.
Methods: A structured comparative evidence synthesis was conducted using peer-reviewed studies. Studies were prioritized when they evaluated multiple detectors on human-authored and AI-generated academic or educational text and reported classification accuracy, false-positive behavior, model-version effects, paraphrasing effects or language-related bias. Findings were synthesized across GPTZero, Turnitin, Copyleaks, Originality.AI, ZeroGPT, Writer, CrossPlag and related detectors. Because platforms, thresholds and test corpora differed substantially and detector versions changed over time, no pooled cross-platform accuracy estimate was calculated.
Results: Detector performance varied markedly by platform, generator model and text transformation. A 2024 benchmark involving 805 detector-text evaluations reported mean accuracy of 39.5% for unmodified AI-generated content across tested detectors, falling to 22.14% after simple adversarial manipulation; human control accuracy was 67%. Earlier multi-tool evaluations similarly found detector accuracy ranging from approximately 43% to 81%, with false-negative rates reaching 100% for some tools and settings. GPT-4 outputs were more difficult to identify than GPT-3.5 outputs. A 2023 fairness study found that 61.3% of TOEFL essays written by non-native English speakers were classified as AI-generated by at least one detector, demonstrating substantial inequity risk. Studies published in 2025 showed that newer detectors can perform strongly on fully machine-generated academic text, yet writing-assistance tools, paraphrasing, hybrid authorship and newer model outputs continued to produce inconsistent classifications.
Conclusion: AI-text detectors can provide useful screening information but do not offer stable authorship attribution across platforms and writing conditions. Detector scores should not be used as standalone evidence of academic misconduct. High-stakes decisions require process evidence, contextual review and human adjudication.

References

Elkhatat AM, Elsaid K, Almeer S. Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text. Int J Educ Integr. 2023;19:17. doi:10.1007/s40979-023-00140-5.

Weber-Wulff D, Anohina-Naumeca A, Bjelobaba S, Foltýnek T, Guerrero-Dib J, Popoola O, et al. Testing of detection tools for AI-generated text. Int J Educ Integr. 2023;19:26. doi:10.1007/s40979-023-00146-z.

Liang W, Yuksekgonul M, Mao Y, Wu E, Zou J. GPT detectors are biased against non-native English writers. Patterns (N Y). 2023;4(7):100779. doi:10.1016/j.patter.2023.100779.

Perkins M, Roe J, Vu BH, Postma D, Hickerson D, McGaughran J, et al. Simple techniques to bypass GenAI text detectors: implications for inclusive education. Int J Educ Technol High Educ. 2024;21:53. doi:10.1186/s41239-024-00487-w.

Bordalejo B, Pafumi D, Onuh F, Khalid AKMI, Pearce MS, O'Donnell DP. Scarlet Cloak and the Forest Adventure: a preliminary study of the impact of AI on commonly used writing tools. Int J Educ Technol High Educ. 2025;22:6. doi:10.1186/s41239-025-00505-5.

Erol G, Ergen A, Gülşen Erol B, Kaya Ergen Ş, Bora TS, Çölgeçen AD, et al. Can we trust academic AI detective? Accuracy and limitations of AI-output detectors. Acta Neurochir (Wien). 2025;167:214. doi:10.1007/s00701-025-06622-4.

Cheng A, et al. Ability of AI detection tools and humans to accurately identify AI-generated content in medical education. BMC Med Educ. 2025. doi:10.1186/s41077-025-00396-6.

Dalalah D, Dalalah OMA. The false positives and false negatives of generative AI detection tools in education and academic research: the case of ChatGPT. Int J Manag Educ. 2023;21(2):100822. doi:10.1016/j.ijme.2023.100822.

Popkov AA, Barrett TS. AI vs academia: experimental study on AI text detectors' accuracy in behavioral health academic writing. Account Res. 2025;32(7):1072-1088. doi:10.1080/08989621.2024.2331757.

Perkins M, Roe J. Academic publisher guidelines on AI usage: a ChatGPT supported thematic analysis. F1000Res. 2024;12:1398.

Cotton DRE, Cotton PA, Shipway JR. Chatting and cheating: ensuring academic integrity in the era of ChatGPT. Innov Educ Teach Int. 2024;61(2):228-239. doi:10.1080/14703297.2023.2190148.

Kasneci E, Sessler K, Küchemann S, Bannert M, Dementieva D, Fischer F, et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learn Individ Differ. 2023;103:102274. doi:10.1016/j.lindif.2023.102274.

Downloads

Published

2026-08-24

Similar Articles

You may also start an advanced similarity search for this article.