Esc
EthicsEmerging

OpenAI Astra ARC-AGI-3 score disputed over custom harness use

Is this a scandal?

Not yet — an early signal. Noise 45/100, holding steady, across 1 source.

SCAND-225554as of Methodology
Cite this incident"OpenAI Astra ARC-AGI-3 score disputed over custom harness use." SCAND.Ai incident SCAND-225554, noise 45/100 as of September 4, 2026. https://scand.ai/scandal/openai-astra-arc-agi-3-benchmark-controversy
FORECASTForecast, not fact

Benchmark organizers will likely mandate standardized harness restrictions for official leaderboards because current variability renders cross-model comparisons scientifically meaningless.

45

Noise 45/100 — louder than 99% of tracked AI controversies.

AI-assisted analysis · How we work

Why it matters

Inconsistent benchmarking methodologies erode trust in AI progress claims and complicate objective comparisons between frontier models.

Key points

  1. OpenAI reported Astra achieving 98.6% on ARC-AGI-3 versus 7.8% for GPT 5.6 Sol using standard harnesses.
  2. Critics note Astra used a custom agentic harness and MAX thinking level unlike competitors' standard evaluations.
  3. Independent repositories show GPT 5.6 Sol and Claude Opus 5 reaching 100% on ARC-AGI-3 with custom harnesses.
  4. ARC-AGI-3 is increasingly viewed as oversaturated and harness-dependent rather than a measure of raw model ability.
  5. The controversy centers on whether technically true metrics were presented without necessary methodological context.

The story

OpenAI faces allegations of misleading benchmark reporting after claiming its Astra model achieved 98.6% on ARC-AGI-3, significantly outperforming competitors. Critics argue this comparison is invalid because Astra utilized a custom agentic harness and maximum thinking settings, while rival scores used standard evaluation protocols. Independent researchers have demonstrated that older models like GPT 5.6 Sol and Claude Opus 5 achieve equal or superior scores when permitted similar custom harness configurations. The controversy highlights growing concerns regarding benchmark saturation and the lack of standardized testing methodologies in frontier AI evaluation. OpenAI has not publicly addressed these specific methodological discrepancies. Industry observers note that without standardized harness restrictions, ARC-AGI-3 may no longer serve as a reliable metric for comparing raw model intelligence. This dispute underscores the urgent need for unified evaluation standards to prevent technically accurate but contextually deceptive performance claims.

Who's involved

Critic
PsychologicalSoup251

Alleges OpenAI deliberately misled the public by comparing Astra's custom-harness score against competitors' standard-harness baselines.

Defender
OpenAI

Reported Astra's 98.6% ARC-AGI-3 score based on internal evaluation methodology without disclosing custom harness usage in primary marketing.

Neutral
ARC Prize Foundation

Maintains public leaderboard data distinguishing between standard and custom evaluation configurations.

How the conversation shifted

the split has narrowed

Polarity (0–100) from the noise pipeline, sampled over time.

Join the Discussion

Discuss this story

Community comments coming in a future update

Be the first to share your perspective. Subscribe to comment.

Noise Level

Buzz45?Noise Score (0–100): how loud a controversy is. Composite of reach, engagement, star power, cross-platform spread, polarity, duration, and industry impact — with 7-day decay.
Decay: 99%
Reach
41
Engagement
84
Star Power
40
Duration
8
Cross-Platform
20
Polarity
50
Industry Impact
50

The timeline

  1. Community verifies alternative scores

    r/singularity users confirm cited repositories demonstrate older models matching Astra under identical conditions.

  2. Reddit analysis challenges methodology

    User PsychologicalSoup251 details custom harness discrepancy and cites GitHub repos showing parity.

  3. OpenAI releases Astra ARC-AGI-3 results

    Company publishes screenshot showing 98.6% score compared to significantly lower competitor baselines.

The full record

Sources & methodology

Every claim above traces to these primary items. How we score →

The forecast

Benchmark organizers will likely mandate standardized harness restrictions for official leaderboards because current variability renders cross-model comparisons scientifically meaningless.

Forecast, not fact — an editorial estimate we score when this resolves.

You're up to date

That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.

Follow this story

We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.

Tracking this story since September 4, 2026.