How does Originality.ai compare to Pangram at identifying AI involvement in academic abstracts?
A recent study, Why AI Detection Fails for Academic Integrity, evaluated the Pangram AI detector. We wanted to see how Originality.ai stacked up.
So, we conducted a study extension to find out. Check out a quick overview below. Then, keep reading for more insight into our findings.
Originality.ai achieved an F1 score of 99.62%, compared with 91.74% for the authors’ released Pangram results, in our analysis of 2,056 academic abstract variants. Every humanized output was excluded.
Originality.ai results were scanned with AI Allowance (15%).
We extended the public dataset accompanying Why AI Detection Fails for Academic Integrity by:
Note: This is an Originality.ai analysis of an external dataset. The authors did not conduct this Originality.ai evaluation, and Pangram was not rerun for this extension.

The source study used 642 English-language abstracts from chemistry, computer science, political science, and theology, spanning 2013–2015 and 2023–2025.
Gemini 3 Flash produced three variants for each paper:
See: Source study.
For our comparison, we retained texts of at least 100 scanner-tokenized words.
This left 2,056 unique texts from 642 papers:
Both detectors are evaluated on exactly these records. All humanized content is excluded from the extension’s results and data workbook.
F1 combines precision and recall. Here, a true positive is an AI-edited or newly generated abstract flagged as AI-involved.
A false positive is an untouched pre-LLM original flagged by the detector.
Originality.ai’s confusion matrix is:
Pangram’s is:
F1 is calculated as 2 × TP / (2 × TP + FP + FN).
Balanced accuracy gives the positive and negative classes equal weight. The sample is 88.91% AI-involved, so its precision, F1, and overall accuracy should not be treated as estimates for a typical classroom or publishing workflow.
The largest detection differences occur in the two editing conditions. These are AI-assisted rewrites:
Neither prompt specifies a percentage of AI-written words or a required Human/Mixed/AI label; See rewriting prompts.
The paper counts LLM rewrites as AI-positive for detection evaluation and discusses an alternative human-source interpretation for abstract-only editing. Our metrics measure AI involvement, not whether assistance violated a rule or exceeded the selected 15% Allowance. Label interpretation.

Excluding both editing conditions leaves 642 newly generated abstracts and the same 228 human controls, or 870 texts.
✓ Originality.ai detected 641 of 642 generated abstracts.
Pangram detected 625.
The F1 difference narrows to 99.38% for Originality.ai versus 98.66% for Pangram.
Balanced accuracy slightly favors Pangram, at 98.68% versus 98.39%, because it produced fewer observed false positives. The metric and intended use both matter when interpreting the comparison.
This study extends and builds on the published Why AI Detection Fails for Academic Integrity study released in August 2026.
The findings of our study extension highlight that Originality.ai performed better (99.62%) than the historical results from Pangram (91.74%) in F1 for AI-involvement detection. As a result, Originality.ai had a 7+ percentage-point lead in that metric.
Originality.ai also identified more AI-involved abstracts. However, there was a slight tradeoff when it came to false positives, of 3.07% for Originality.ai vs Pangram’s 0%.
Interested in learning more about Originality.ai compared to Pangram? Check out our in-depth Originality.ai vs Pangram Study.
Further Reading:
We used the authors’ released repository, frozen at commit 589525a. Originality.ai was scanned on 21 September 2026 with 15% AI Allowance through API v3. The decision rule was AI confidence ≥0.50. Pangram used the authors’ saved AI-likelihood scores at the same numerical cutoff. The stored Pangram texts matched the corresponding Originality.ai inputs exactly.
This extension was requested after the broader results were available. It reuses the frozen sample and saved scores without changing thresholds or selecting records by detector outcome. Independent calculations reproduced the confusion matrices and metrics. Restricting the analysis to the earlier source period produced F1 scores of 99.23% for Originality.ai and 90.21% for Pangram.
The comparison has four important limits:
Historical competitor results: Pangram was not tested contemporaneously. Its saved pre-humanization files lack a model identifier; the Originality.ai API also returned no exact build identifier.
Assisted writing is a distinct question: The corpus does not establish exact AI contribution percentages, so these results cannot validate 15% Allowance compliance or determine misconduct.
Limited scope: Four academic domains and one rewrite generator do not represent all writing. Several variants come from each paper, so observations are dependent. We report descriptive point estimates rather than statistical-superiority claims.
Observed error rates have limits: Zero Pangram false positives among 228 controls does not establish zero false positives in general. Excluding humanized content means this extension makes no claim about humanizer resistance.
Originality.ai achieved higher F1 and recall on this defined AI-involvement task. Pangram produced fewer false positives. These findings support comparing the metrics against the intended use, with human review where authorship or academic integrity decisions are involved.
