AI Studies

Is GPT-6 Astra Content Detectable?

We tested 1,000 English GPT-6 Astra responses with Originality.ai AI detector to find out if GPT-6 Astra content was detectable. We also compared length and readability with Kimi K3 and archived responses from OpenAI, Anthropic, DeepSeek, Google and xAI on the same prompts.

We tested 1,000 English GPT-6 Astra responses with Originality.ai AI detector to find out if GPT-6 Astra content was detectable. We also compared length and readability with Kimi K3 and archived responses from OpenAI, Anthropic, DeepSeek, Google and xAI on the same prompts.

The quick answer

Yes, in this benchmark. Originality.ai AI Detection classified 999 of 1,000 GPT-6 Astra outputs as Likely AI, a 99.9% detection rate.

Three findings stood out:

  • GPT-6 Astra was detectable: 999/1,000 final outputs reached the 50% AI-score threshold.
  • GPT-6 wrote less than Kimi K3 on identical prompts: it was shorter in 681/924 comparisons (73.7%), but longer than the matched historical response in 824/924 (89.2%).
  • Shorter did not necessarily mean easier to read: Originality.ai gave GPT-6 a median paired Flesch-Kincaid grade 1.1 levels above Kimi K3.
The quick answer

How we tested GPT-6 Astra

This study follows the writing-task format used in previous tests.

Archived comparison models

The archived responses came from 17 model labels. Each prompt has one archived response, not a response from every model:

  • OpenAI: ChatGPT-4o latest, GPT-4o, GPT-4o mini, GPT-4.1, GPT-4.1 mini and o4-mini.
  • Anthropic: Claude 3.5 Haiku, Claude 3.7 Sonnet, Claude Opus 4 and Claude Sonnet 4.
  • DeepSeek: DeepSeek Chat and DeepSeek Reasoner.
  • Google: Gemini 2.0 Flash, Gemini 2.0 Flash-Lite and Gemini 2.5 Pro Preview.
  • xAI: Grok 3 and Grok 3 mini.

The sample contained 200 prompts in each of five content types: blogs, conversational writing, expository writing, news and reviews. Some briefs asked for human-like style or wording intended to be difficult for detectors to distinguish; these were retained, not rewritten for this test.

GPT-6 generation and detection took place on Sep 4, 2026. We used the OpenAI Responses API with the model ID gpt-6-astra. 

Only AI Allowance 15% results are included here; a score of at least 0.50 counts as detected. The 15% setting is a detector mode, not the percentage of words we claim were written by AI.

Content type Detected / tested Detection rate
Blog 200 / 200 100%
Conversational 199 / 200 99.5%
Expository 200 / 200 100%
News 200 / 200 100%
Review 200 / 200 100%
Total 999 / 1,000 99.9%

The one below-threshold response was conversational.

So GPT-6 Astra is detectable but what else can we understand about the text it produces compared to previous models? 

How long were the responses on the same prompts?

Across the same prompts, the median GPT-6 response was 892 words, compared with 1,112.5 for Kimi K3 and 269.5 for the archived responses. GPT-6 was shorter than Kimi in 681/924 pairs (73.7%); the median within-prompt difference was 225.5 fewer words. Against the historical response, GPT-6 was longer in 824/924 pairs (89.2%), with a median paired difference of 422 more words.

How long were the responses on the same prompts?

Length depended on the assignment. Expository writing produced the longest GPT-6 responses; news produced the shortest. Blog medians were similar, even though GPT-6 was shorter overall.

Prompt category Matched prompts GPT-6 median words Kimi K3 median words Historical median words
Blog 199 1,144 1,127 546
Conversational 195 722 868 215
Expository 177 1,442 1,701 1,233
News 156 394.5 889.5 158.5
Review 197 875 1,194 194

Did GPT-6 follow requested word counts?

Only two of the 1,000 prompts contained explicit numeric word-count targets. Both were approximate targets, and both remained unchanged. Most prompts specified the subject, style or format instead, so the dataset cannot establish a general word-count compliance rate.

Approximate target GPT-6 Astra Kimi K3 Archived response
800 words 855 (+6.9%) 811 (+1.4%) Grok 3: 583 (−27.1%)
3,000 words 2,981 (−0.6%) 2,941 (−2.0%) GPT-4.1 mini: 1,809 (−39.7%)

GPT-6 and Kimi came close to the requested lengths in these two examples. Two examples are not enough to rank instruction-following ability.

How readable was GPT-6's writing?

We used the Originality.ai Readability Checker on the complete responses to all 924 unchanged prompts. Its Flesch-Kincaid grade level estimates reading difficulty: lower scores mean simpler text.

Median grades were 11.0 for GPT-6, 8.9 for Kimi K3 and 10.2 for the archived responses. Comparing each prompt individually, GPT-6’s median gap was +1.1 grades versus Kimi and 0.0 versus the archive.

How readable was GPT-6's writing?

These scores estimate text complexity, not factual accuracy, writing quality or reader comprehension.

Methodology and limitations

  • Sample: 1,000 English GPT-6 Astra responses; 200 in each of five writing categories.
  • Detection: AI Allowance 15%; scores of 0.50 or higher counted as detected, matched to the exact response text.
  • Final outputs: The latest successful response per prompt. All 1,000 final scans succeeded.
  • Comparisons: Length and readability use 924 unchanged prompts, some prompts failed initially to produce a 100 word response so were adjusted. These adjusted prompts and responses were excluded from the study.
  • Readability: Originality.ai Readability Checker, Flesch-Kincaid grade levels.
  • Limits: English, known-AI text only; no human control group or false-positive measurement. Length and readability comparisons are exploratory, not assessments of writing quality.

Final thoughts

Originality.ai AI Allowance 15% identified 999 of 1,000 final GPT-6 Astra outputs, a 99.9% detection rate in this test.

Jonathan Gillham

Jonathan Gillham

Jonathan Gillham is an engineer, inventor, and entrepreneur. He is the founder and CEO of Originality.ai, an AI content integrity platform that launched the first commercial AI detector in November 2022, just three days before ChatGPT launched. Before founding Originality.ai, Jon worked as an engineer, built and exited two companies. His early work with generative AI in 2020 and 2021 gave him a firsthand view of the coming wave of AI-generated content and the need for technology that could bring transparency and trust to written content. Today, he leads Originality.ai’s work in AI detection and content integrity and is a named inventor on two U.S. patents covering AI detection technology. Jon’s expertise and research have been featured in WIRED, Business Insider, The Register, Global News, The Guardian, Entrepreneur, and The Washington Post, among others.

Al Content Detector & Plagiarism Checker for Marketers and Writers

Use our leading tools to ensure you can hit publish with integrity!

Try our AI Checker now!

cross image
Free Tool Popup image

Sign up now!

Free Tool Image step1
Free Tool Image step2
Free Tool Image step3
Free Tool Image step1
Free Tool Image step2
Free Tool Image step3
Free Tool Image step4
Free Tool Image step5