We tested 1,000 English GPT-6 Astra responses with Originality.ai AI detector to find out if GPT-6 Astra content was detectable. We also compared length and readability with Kimi K3 and archived responses from OpenAI, Anthropic, DeepSeek, Google and xAI on the same prompts.
Yes, in this benchmark. Originality.ai AI Detection classified 999 of 1,000 GPT-6 Astra outputs as Likely AI, a 99.9% detection rate.
Three findings stood out:

This study follows the writing-task format used in previous tests.
The archived responses came from 17 model labels. Each prompt has one archived response, not a response from every model:
The sample contained 200 prompts in each of five content types: blogs, conversational writing, expository writing, news and reviews. Some briefs asked for human-like style or wording intended to be difficult for detectors to distinguish; these were retained, not rewritten for this test.
GPT-6 generation and detection took place on Sep 4, 2026. We used the OpenAI Responses API with the model ID gpt-6-astra.
Only AI Allowance 15% results are included here; a score of at least 0.50 counts as detected. The 15% setting is a detector mode, not the percentage of words we claim were written by AI.
The one below-threshold response was conversational.
So GPT-6 Astra is detectable but what else can we understand about the text it produces compared to previous models?
Across the same prompts, the median GPT-6 response was 892 words, compared with 1,112.5 for Kimi K3 and 269.5 for the archived responses. GPT-6 was shorter than Kimi in 681/924 pairs (73.7%); the median within-prompt difference was 225.5 fewer words. Against the historical response, GPT-6 was longer in 824/924 pairs (89.2%), with a median paired difference of 422 more words.

Length depended on the assignment. Expository writing produced the longest GPT-6 responses; news produced the shortest. Blog medians were similar, even though GPT-6 was shorter overall.
Only two of the 1,000 prompts contained explicit numeric word-count targets. Both were approximate targets, and both remained unchanged. Most prompts specified the subject, style or format instead, so the dataset cannot establish a general word-count compliance rate.
GPT-6 and Kimi came close to the requested lengths in these two examples. Two examples are not enough to rank instruction-following ability.
We used the Originality.ai Readability Checker on the complete responses to all 924 unchanged prompts. Its Flesch-Kincaid grade level estimates reading difficulty: lower scores mean simpler text.
Median grades were 11.0 for GPT-6, 8.9 for Kimi K3 and 10.2 for the archived responses. Comparing each prompt individually, GPT-6’s median gap was +1.1 grades versus Kimi and 0.0 versus the archive.

These scores estimate text complexity, not factual accuracy, writing quality or reader comprehension.
Originality.ai AI Allowance 15% identified 999 of 1,000 final GPT-6 Astra outputs, a 99.9% detection rate in this test.
