AI Studies

Study Extension: Comparing Originality.ai and Pangram on Academic Abstracts

Find out how Originality.ai compares to Pangram in identifying AI involvement in academic abstracts in this study extension.

How does Originality.ai compare to Pangram at identifying AI involvement in academic abstracts?

A recent study, Why AI Detection Fails for Academic Integrity, evaluated the Pangram AI detector. We wanted to see how Originality.ai stacked up. 

So, we conducted a study extension to find out. Check out a quick overview below. Then, keep reading for more insight into our findings.

The quick results:

Originality.ai achieved an F1 score of 99.62%, compared with 91.74% for the authors’ released Pangram results, in our analysis of 2,056 academic abstract variants. Every humanized output was excluded.

Originality.ai results were scanned with AI Allowance (15%). 

How did we extend the study?

We extended the public dataset accompanying Why AI Detection Fails for Academic Integrity by:

  • Adding Originality.ai scans
  • Calculating comparable metrics on the same eligible texts

Note: This is an Originality.ai analysis of an external dataset. The authors did not conduct this Originality.ai evaluation, and Pangram was not rerun for this extension.

Key Study Extension Findings:

  • Originality.ai Had Higher F1 for AI-involvement detection: 99.62% for Originality.ai versus 91.74% for historical Pangram, a 7.88 percentage-point difference.
  • Originality.ai Detected More AI-Involved Abstracts: Originality.ai flagged 1,821 of 1,828 edited or generated abstracts (99.62%); Pangram flagged 1,549 (84.74%).
  • The false-positive trade-off: Originality.ai flagged 7 of 228 untouched pre-LLM originals (3.07%). Pangram flagged 0 of 228 (0.00%).
F1 for Detecting AI Involvement, Study Extension

Dataset and Study Scope

The source study used 642 English-language abstracts from chemistry, computer science, political science, and theology, spanning 2013–2015 and 2023–2025. 

Gemini 3 Flash produced three variants for each paper: 

  • an abstract-only rewrite
  • a rewrite informed by the full article
  • a new abstract from the article. 

See: Source study.

For our comparison, we retained texts of at least 100 scanner-tokenized words. 

  • Of 2,568 records before humanization, 196 fell below that minimum. 
  • We also excluded 316 eligible recent originals because their underlying AI use is unknown. 

This left 2,056 unique texts from 642 papers:

  • 1,828 positives: AI-edited or newly generated abstracts from both source periods.
  • 228 negative controls: untouched original abstracts published in 2013–2015, used as a proxy-human reference.

Both detectors are evaluated on exactly these records. All humanized content is excluded from the extension’s results and data workbook.

F1 and Other Detection Metrics

F1 combines precision and recall. Here, a true positive is an AI-edited or newly generated abstract flagged as AI-involved. 

A false positive is an untouched pre-LLM original flagged by the detector.

Metric Originality.ai 15% AI Allowance Pangram historical
F1 score 99.62% 91.74%
Precision 99.62% 100.00%
Recall 99.62% 84.74%
Accuracy 99.32% 86.43%
False-positive rate 3.07% 0.00%
False-negative rate 0.38% 15.26%
Specificity 96.93% 100.00%
Balanced accuracy 98.27% 92.37%

‍

Originality.ai’s confusion matrix is:

  • 1,821 true positives
  • 7 false positives
  • 221 true negatives
  • 7 false negatives

Pangram’s is:

  • 1,549 true positives
  • 0 false positives
  • 228 true negatives
  • 279 false negatives

F1 is calculated as 2 × TP / (2 × TP + FP + FN).

Balanced accuracy gives the positive and negative classes equal weight. The sample is 88.91% AI-involved, so its precision, F1, and overall accuracy should not be treated as estimates for a typical classroom or publishing workflow.

Where the Results Differ

Text condition Eligible texts Originality.ai flagged Pangram flagged
Abstract-only edits 556 552 (99.28%) 409 (73.56%)
Article-informed edits 630 628 (99.68%) 515 (81.75%)
Newly generated abstracts 642 641 (99.84%) 625 (97.35%)
Untouched pre-LLM originals 228 7 (3.07%) 0 (0.00%)

‍

The largest detection differences occur in the two editing conditions. These are AI-assisted rewrites: 

  • Abstract-only editing uses the supplied abstract
  • Article-informed editing also receives the full article 

Neither prompt specifies a percentage of AI-written words or a required Human/Mixed/AI label; See rewriting prompts.

The paper counts LLM rewrites as AI-positive for detection evaluation and discusses an alternative human-source interpretation for abstract-only editing. Our metrics measure AI involvement, not whether assistance violated a rule or exceeded the selected 15% Allowance. Label interpretation.

Detection Rewrite by Condition

Generation Only Comparison

Excluding both editing conditions leaves 642 newly generated abstracts and the same 228 human controls, or 870 texts. 

✓ Originality.ai detected 641 of 642 generated abstracts.

Pangram detected 625.

The F1 difference narrows to 99.38% for Originality.ai versus 98.66% for Pangram. 

Balanced accuracy slightly favors Pangram, at 98.68% versus 98.39%, because it produced fewer observed false positives. The metric and intended use both matter when interpreting the comparison.

Final Thoughts

This study extends and builds on the published Why AI Detection Fails for Academic Integrity study released in August 2026.

The findings of our study extension highlight that Originality.ai performed better (99.62%) than the historical results from Pangram (91.74%) in F1 for AI-involvement detection. As a result, Originality.ai had a 7+ percentage-point lead in that metric.

Originality.ai also identified more AI-involved abstracts. However, there was a slight tradeoff when it came to false positives, of 3.07% for Originality.ai vs Pangram’s 0%.

Interested in learning more about Originality.ai compared to Pangram? Check out our in-depth Originality.ai vs Pangram Study.

Further Reading: 

Methodology and Limitations

We used the authors’ released repository, frozen at commit 589525a. Originality.ai was scanned on 21 September 2026 with 15% AI Allowance through API v3. The decision rule was AI confidence ≥0.50. Pangram used the authors’ saved AI-likelihood scores at the same numerical cutoff. The stored Pangram texts matched the corresponding Originality.ai inputs exactly.

This extension was requested after the broader results were available. It reuses the frozen sample and saved scores without changing thresholds or selecting records by detector outcome. Independent calculations reproduced the confusion matrices and metrics. Restricting the analysis to the earlier source period produced F1 scores of 99.23% for Originality.ai and 90.21% for Pangram.

The comparison has four important limits:

Historical competitor results: Pangram was not tested contemporaneously. Its saved pre-humanization files lack a model identifier; the Originality.ai API also returned no exact build identifier.

Assisted writing is a distinct question: The corpus does not establish exact AI contribution percentages, so these results cannot validate 15% Allowance compliance or determine misconduct.

Limited scope: Four academic domains and one rewrite generator do not represent all writing. Several variants come from each paper, so observations are dependent. We report descriptive point estimates rather than statistical-superiority claims.

Observed error rates have limits: Zero Pangram false positives among 228 controls does not establish zero false positives in general. Excluding humanized content means this extension makes no claim about humanizer resistance.

Originality.ai achieved higher F1 and recall on this defined AI-involvement task. Pangram produced fewer false positives. These findings support comparing the metrics against the intended use, with human review where authorship or academic integrity decisions are involved.

Jonathan Gillham

Jonathan Gillham

Jonathan Gillham is an engineer, inventor, and entrepreneur. He is the founder and CEO of Originality.ai, an AI content integrity platform that launched the first commercial AI detector in November 2022, just three days before ChatGPT launched. Before founding Originality.ai, Jon worked as an engineer, built and exited two companies. His early work with generative AI in 2020 and 2021 gave him a firsthand view of the coming wave of AI-generated content and the need for technology that could bring transparency and trust to written content. Today, he leads Originality.ai’s work in AI detection and content integrity and is a named inventor on two U.S. patents covering AI detection technology. Jon’s expertise and research have been featured in WIRED, Business Insider, The Register, Global News, The Guardian, Entrepreneur, and The Washington Post, among others.

Al Content Detector & Plagiarism Checker for Marketers and Writers

Use our leading tools to ensure you can hit publish with integrity!

Try our AI Checker now!

cross image
Free Tool Popup image

Sign up now!

Free Tool Image step1
Free Tool Image step2
Free Tool Image step3
Free Tool Image step1
Free Tool Image step2
Free Tool Image step3
Free Tool Image step4
Free Tool Image step5