2024 Study on Large Language Model-Generated Text in Medical Writing Concludes That Originality.ai Excels at Detecting AI-Generated and AI-Paraphrased Text
Originality.ai is exceptional at identifying AI-generated medical articles, according to the study “The Great Detectives: Humans vs. AI Detectors in Catching Large Language Model-Generated Medical Writing,” 2024.
A recent study (published in 2024) in the International Journal for Educational Integrity, “The Great Detectives: Humans vs. AI Detectors in Catching Large Language Model-Generated Medical Writing,” compared the efficacy of leading AI content detectors to human reviewers.
The study concluded that Originality.ai was the most effective AI detector for identifying AI-generated rehabilitation-related articles in medical writing — in their original and paraphrased versions.
Key Findings (TL;DR)
Originality.ai emerged as the most sensitive and accurate platform for detecting 100% of ChatGPT-generated and paraphrased content.
In contrast, human reviewers were less accurate. Student reviewers only identified 76% of AI-paraphrased articles correctly. Then, professors who reviewed the articles misclassified 12% of human-written content as AI-generated.
Study Details
The study evaluated six common AI content detectors and four human reviewers (two students and two professors). Then, both the AI-generated text detector and the human reviewers had the task of distinguishing between 150 academic papers consisting of original, ChatGPT-generated, and AI-rephrased content.
(An outline of the study)
AI Text Detection Tools
Six AI Text Detectors Used: Originality.ai, TurnItIn, GPTZero, ZeroGPT, Content at Scale, and GPT-2 Output Detector.
Four Human Reviewers: Two students and two professors.
Dataset
The dataset consisted of 150 academic papers.
50 Original Papers: Rehabilitation-related articles from four peer-reviewed journals.
50 AI-Generated Papers: ChatGPT generated the introduction, discussion, and conclusion sections based on the original titles, methods, and results.
50 AI-Rephrased Papers: Wordtune rephrased the ChatGPT-generated articles.
Evaluation Criteria
Accuracy, misclassification rate, ROC curve, sensitivity, specificity, time taken (only for human reviewers), and reasons for classification (only for human reviewers).
Originality.ai’s AI Detector Results
Finding 1: Originality.ai detected 100% of both ChatGPT-generated and AI-rephrased articles
(The accuracy of six AI content detectors in identifying AI-generated articles)
Finding 2: Originality.ai scored the highest mean AI Score
For ChatGPT-generated articles: 98.74%
For AI-rephrased articles: 99.74%
(The mean AI scores of 50 ChatGPT-generated articles before and after rephrasing)
Finding 3: Originality.ai ranked third for the lowest percentage of misclassification or uncertainty
(the percentage of misclassification of human-written articles as AI-generated ones by detectors)
Turnitin and GPT-2 Output Detector were excellent at identifying human-written texts but not as effective with AI-rephrased content.
Finding 4: Human reviewers were less accurate
Student reviewers identified only 76% of AI-rephrased articles.
Professors misclassified 12% of human-written texts as AI-generated.
In the graphs below, Reviewer 1 and Reviewer 2 are college students, while Reviewer 3 and Reviewer 4 are professors.
(the accuracy of four human reviewers in identifying AI-rephrased articles)
(the percentage of misclassifying human-written articles as AI-rephrased ones by reviewers)
Final Thoughts
Of the AI detectors and human reviewers involved in the study, Originality.ai stands out as the most effective AI-generated text detection tool for identifying AI-generated medical writing (including paraphrased content). Using Originality.ai can significantly enhance the peer-review process and uphold academic integrity in scientific publishing.