OpenAI Astra ARC-AGI-3 score disputed over custom harness use
Is this a scandal?
Not yet — an early signal. Noise 45/100, holding steady, across 1 source.
Benchmark organizers will likely mandate standardized harness restrictions for official leaderboards because current variability renders cross-model comparisons scientifically meaningless.
Noise 45/100 — louder than 99% of tracked AI controversies.
Why it matters
Inconsistent benchmarking methodologies erode trust in AI progress claims and complicate objective comparisons between frontier models.
Key points
- OpenAI reported Astra achieving 98.6% on ARC-AGI-3 versus 7.8% for GPT 5.6 Sol using standard harnesses.
- Critics note Astra used a custom agentic harness and MAX thinking level unlike competitors' standard evaluations.
- Independent repositories show GPT 5.6 Sol and Claude Opus 5 reaching 100% on ARC-AGI-3 with custom harnesses.
- ARC-AGI-3 is increasingly viewed as oversaturated and harness-dependent rather than a measure of raw model ability.
- The controversy centers on whether technically true metrics were presented without necessary methodological context.
The story
OpenAI faces allegations of misleading benchmark reporting after claiming its Astra model achieved 98.6% on ARC-AGI-3, significantly outperforming competitors. Critics argue this comparison is invalid because Astra utilized a custom agentic harness and maximum thinking settings, while rival scores used standard evaluation protocols. Independent researchers have demonstrated that older models like GPT 5.6 Sol and Claude Opus 5 achieve equal or superior scores when permitted similar custom harness configurations. The controversy highlights growing concerns regarding benchmark saturation and the lack of standardized testing methodologies in frontier AI evaluation. OpenAI has not publicly addressed these specific methodological discrepancies. Industry observers note that without standardized harness restrictions, ARC-AGI-3 may no longer serve as a reliable metric for comparing raw model intelligence. This dispute underscores the urgent need for unified evaluation standards to prevent technically accurate but contextually deceptive performance claims.
Who's involved
Alleges OpenAI deliberately misled the public by comparing Astra's custom-harness score against competitors' standard-harness baselines.
Reported Astra's 98.6% ARC-AGI-3 score based on internal evaluation methodology without disclosing custom harness usage in primary marketing.
Maintains public leaderboard data distinguishing between standard and custom evaluation configurations.
How the conversation shifted
Polarity (0–100) from the noise pipeline, sampled over time.
Noise Level
The timeline
Community verifies alternative scores
r/singularity users confirm cited repositories demonstrate older models matching Astra under identical conditions.
Reddit analysis challenges methodology
User PsychologicalSoup251 details custom harness discrepancy and cites GitHub repos showing parity.
OpenAI releases Astra ARC-AGI-3 results
Company publishes screenshot showing 98.6% score compared to significantly lower competitor baselines.
The full record
Sources & methodology
Every claim above traces to these primary items. How we score →
The forecast
Benchmark organizers will likely mandate standardized harness restrictions for official leaderboards because current variability renders cross-model comparisons scientifically meaningless.
Forecast, not fact — an editorial estimate we score when this resolves.
That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.
Follow this story
We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.
Tracking this story since September 4, 2026.
Join the Discussion
Discuss this story
Community comments coming in a future update
Be the first to share your perspective. Subscribe to comment.