ThinkPatternGet the app
Story
TECHNOLOGY · SEP 5, 2026

OpenAI Modifies GPT-6 Astra Benchmarks After Retracted Blog Post

OpenAI updated performance metrics for its GPT-6 Astra model following a troubled announcement rollout, sparking accusations of benchmaxxing from Stanford researchers.

OpenAI modified several evaluation benchmarks for its GPT-6 Astra model after a problematic blog post rollout on September 3. The company initially published the announcement at 2 p.m. ET but retracted it for undisclosed reasons, later republishing the post with updated metrics. Some of these changes favored Astra or disadvantaged competitors, such as Anthropic.

Specific fluctuations included Astra's hallucination rate, which was briefly halved from 4.2% to 2% before returning to the original figure. The company also adjusted mathematics scores for its own GPT-5.6 Sol and Anthropic's Fable 5.1. OpenAI stated these adjustments were necessary to ensure the numbers represented the best estimate of model performance.

Researchers from the Stanford Intelligent Systems Laboratory characterized the changes as potential benchmaxxing, a practice of manipulating testing conditions to maximize scores for marketing purposes. The incident highlights broader industry struggles with the standardization and transparency of large language model benchmarks.


Reported across 2 outlets
Actors
OpenAIAnthropicStanford Intelligent Systems Laboratory

Keep reading in the app

The full story and every source, free in the app.

Download on the App StoreComing soonGoogle Play