Esc
SafetyEmerging

Astra and Fable criticized for reusing outdated alignment evals

Is this a scandal?

Not yet — an early signal. Noise 49/100, heating up, across 2 sources.

SCAND-239782as of Methodology
Cite this incident"Astra and Fable criticized for reusing outdated alignment evals." SCAND.Ai incident SCAND-239782, noise 49/100 as of September 14, 2026. https://scand.ai/scandal/astra-fable-criticized-outdated-alignment-evals
FORECASTForecast, not fact

Expect Astra and Fable to publish updated evaluation methodologies within weeks because silence on safety metrics damages credibility with enterprise customers and regulators.

Confidence: A close call (~60%)

Next to watch: Publication of a technical blog post or arXiv paper by either lab's safety team addressing evaluation methodologies.

How we reached this call
49

Noise 49/100 — louder than 99% of tracked AI controversies.

AI-assisted analysis · How we work

Why it matters

Recycling outdated safety tests risks deploying models with unmeasured dangerous capabilities, undermining industry trust in alignment claims.

Key points

  1. Critics allege Astra and Fable use simple variants of 2025 alignment evaluations for current models.
  2. Outdated benchmarks may fail to detect dangerous capabilities in newer frontier systems.
  3. Neither Astra nor Fable has publicly responded to the evaluation methodology criticism.
  4. The dispute reflects tension between rapid deployment cycles and evolving safety science.
  5. Industry consensus holds that alignment evaluations require continuous updating to remain valid.

The story

AI safety researchers have accused Astra and Fable of continuing to evaluate model alignment using simple variants of benchmarks established in 2025. Critics argue these outdated evaluations fail to capture emerging risks in current frontier models, potentially masking dangerous capabilities. The allegation suggests both companies may be prioritizing development speed over rigorous safety validation by relying on legacy testing frameworks. Neither Astra nor Fable has publicly addressed the specific criticism regarding their evaluation methodologies. Industry observers note that alignment science evolves rapidly, making year-old benchmarks insufficient for assessing modern systems. This dispute highlights growing tensions between rapid AI deployment and the need for updated safety standards. The controversy underscores broader concerns about whether current industry practices adequately measure alignment as model capabilities advance beyond original test parameters.

Who's involved

Critic
AI Safety Researchers

Astra and Fable rely on obsolete 2025 benchmarks that inadequately assess current model risks.

Defender
Astra

Has not publicly responded to allegations of using outdated alignment evaluation methods.

Defender
Fable

Has not publicly responded to allegations of using outdated alignment evaluation methods.

Most contested claim

Astra and Fable are actively hacking on simple variants of 2025 alignment evals

Biggest open question

It is unverified whether Astra and Fable have actually remained silent or if responses exist outside the provided source set

Read the full story

How we got here

The tension between rapid model scaling and evaluation latency is a persistent structural pattern in AI safety research. Historically, benchmark saturation precedes methodological crises, where static tests cease to discriminate between capable models and instead measure test-set memorization or proxy optimization. This cycle has repeated across domains from natural language understanding to code generation, typically triggering a 'benchmark retirement' phase followed by the introduction of dynamic or adversarial evaluation frameworks. The current dispute mirrors prior controversies where leaderboard dominance was decoupled from real-world utility due to eval drift. In alignment specifically, the transition from behavioral compliance checks to mechanistic interpretability or adversarial robustness testing often lags behind capability jumps by 12-18 months. This lag creates windows where safety claims rely on proxies that researchers consider insufficient for the new capability frontier. The pattern suggests that evaluation obsolescence is not merely a company-specific failure but an endemic feature of exponential capability growth outpacing measurement science.

The full story

On September 13, 2026, a controversy emerged within the AI safety and evaluation community regarding the alignment testing methodologies employed by Astra and Fable. The dispute centers on allegations that both companies continue to rely on outdated evaluation frameworks originally established in early 2025, despite significant advancements in model capabilities since that time. According to a post on Hacker News by user Levitating, both Astra and Fable are accused of 'hacking on simple variants' of these older benchmarks rather than developing robust assessments for current-generation risks [1]. This criticism suggests that the safety evaluations currently used to validate models like GPT-6 Astra and Claude-5-Fable may be functionally obsolete, potentially creating a false sense of security regarding their alignment.

The core allegation is that the benchmarks in question were designed for an earlier generation of models and fail to capture the nuanced failure modes or dangerous capabilities of systems deployed in late 2026. User Levitating’s critique draws a direct parallel to 'cheating in chess,' implying that optimizing for these static tests does not correlate with genuine safety or alignment but rather represents a form of metric gaming [1]. This perspective posits that high scores on these legacy evals are artifacts of overfitting rather than evidence of reliable behavior. The timing of this critique coincides with broader industry turbulence regarding benchmark validity; separate reports indicate that Artificial Analysis recently replaced its own intelligence index after discovering similar optimization issues, suggesting a systemic challenge in maintaining evaluation integrity as models improve rapidly [2].

Despite the specificity of these allegations, neither Astra nor Fable has issued a public response as of the latest available information. Their silence leaves several key factual questions unresolved, including whether they have internally updated their evaluation suites without public disclosure, or if they genuinely rely on the 2025 variants criticized by researchers. The lack of rebuttal also makes it difficult to determine if the companies dispute the characterization of their methods as 'hacking' or if they acknowledge the limitations but face technical constraints in developing next-generation evals. In the absence of official statements, the narrative remains driven entirely by external critics and independent observers who argue that the gap between benchmark performance and actual safety assurance has widened dangerously.

The controversy highlights a recurring tension in AI development: the lag between model capability and evaluation methodology. Critics argue that reusing 2025-era tests for 2026-era models is akin to using a high school math test to evaluate a PhD candidate; it measures basic competence but fails to assess advanced reasoning or novel risks. The comparison to chess cheating underscores the adversarial nature of modern evaluation, where models (or their developers) may exploit known patterns in test sets to achieve superficially impressive results. While the specific technical deficiencies of the 2025 benchmarks are not detailed in the available sources, the implication is that they lack the dynamic or adversarial components necessary to stress-test contemporary foundation models effectively.

This incident occurs against a backdrop of increasing skepticism toward proprietary benchmarks. The recent adjustment of the Artificial Analysis Intelligence Index, which saw leadership changes after replacing the τ³ benchmark, demonstrates how volatile performance rankings can be when evaluation criteria shift [2]. Such volatility reinforces critic arguments that static benchmarks are insufficient for tracking true progress or safety. For Astra and Fable, the stakes involve not just reputational damage but the fundamental credibility of their safety claims. If their alignment assurances rest on tests that the research community deems obsolete, stakeholders may question the reliability of their deployments until more rigorous validation methods are demonstrated or disclosed.

What's confirmed, what's disputed

  • ConfirmedUser Levitating alleges Astra and Fable still hack on simple variants of alignment evals from 2025
  • ConfirmedCritics compare the reuse of 2025 alignment evals to cheating in chess
  • ConfirmedArtificial Analysis replaced its τ³ benchmark with a new private eval in the Intelligence Index v4.3 update
  • DisputedAstra and Fable have not publicly responded to allegations of using outdated alignment evaluation methods
  • ConfirmedGPT-6 Astra and Claude-5-Fable are the specific models implicated in the outdated eval controversy

The strongest case each way

Critic's case

Reusing 2025-era alignment benchmarks for 2026-class models constitutes a form of metric hacking analogous to chess cheating, where optimization targets static test patterns rather than genuine safety properties, thereby producing misleading assurance

Defender's case

No public defense is available in the provided sources; however, a potential steelman would be that continuity in evaluation enables longitudinal safety tracking and that 2025 baselines remain necessary regression tests even if supplemented by newer methods

Times this happened before

  • MMLU Saturation Crisis · 2024Community shifted to MMLU-Pro and GPQA; static benchmarks widely deprecated for frontier evaluation
  • HumanEval Contamination Scandal · 2024Code generation leaderboards reset; live coding evals adopted

What's at stake

AI safety researchers lose confidence in published alignment metrics, complicating peer review and collaborative safety standards. Downstream enterprise customers deploying Astra or Fable models may operate under false safety assumptions if 2025 evals miss 2026-era failure modes. The broader ecosystem faces increased uncertainty about which safety claims are trustworthy, potentially slowing adoption or triggering premature regulatory intervention based on incomplete evidence. Magnitude is currently limited to reputational and epistemic harm; no financial penalties, user counts, or revenue figures are cited in available sources. The primary risk is cascading loss of trust in alignment evaluation as a discipline if multiple leading labs are simultaneously found to be gaming obsolete benchmarks.

2025 alignment eval variants vs 2026 modelsBenchmark versions at issue
2 revisionsIntelligence Index updates in 3 days

What we still don't know

  • It is unverified whether Astra and Fable have actually remained silent or if responses exist outside the provided source set

How the conversation shifted

the split has narrowed

Polarity (0–100) from the noise pipeline, sampled over time.

Join the Discussion

Discuss this story

Community comments coming in a future update

Be the first to share your perspective. Subscribe to comment.

Noise Level

Buzz49?Noise Score (0–100): how loud a controversy is. Composite of reach, engagement, star power, cross-platform spread, polarity, duration, and industry impact — with 7-day decay.
Decay: 99%
Reach
48
Engagement
81
Star Power
30
Duration
20
Cross-Platform
50
Polarity
50
Industry Impact
50

The timeline

  1. Hacker News post criticizes Astra and Fable

    User Levitating alleges both companies still hack on simple variants of 2025 evals.

  2. Original alignment benchmarks established

    Baseline evaluation frameworks created that critics now claim are obsolete.

The full record

Sources & methodology

Every claim above traces to these primary items. How we score →

Where the sources disagree

In dispute Astra and Fable are actively hacking on simple variants of 2025 alignment evals

Established A Hacker News user named Levitating has publicly alleged that Astra and Fable rely on outdated 2025 alignment eval variants, comparing this practice to cheating; no public response from either company is documented in available sources

What's being under-reported

Under-reported by mainstream

Heavily discussed on social platforms, but not yet covered by any news outlet.

  • Coverage: 4 social posts, 0 news-outlet items.
  • Voices: 1 critic, 2 defenders.

Coverage lacks direct input from Astra and Fable engineering teams, creating a one-sided narrative driven entirely by external critics. Also missing are perspectives from enterprise customers who rely on these alignment claims for procurement decisions; their risk tolerance and verification practices would contextualize whether this controversy has real-world deployment consequences or remains an academic dispute. Without defender voices, the dossier cannot assess whether the criticism is technically accurate or rhetorically amplified.

Who changed their mind, and why
  • AI Safety ResearchersEscalated from general benchmark skepticism to naming Astra and Fable specifically as offenders relying on 2025 eval variants (was: Broad concern about evaluation-capability gap)
  • AstraNo observable position change; silence maintained through September 13, 2026 (was: Unknown)
  • FableNo observable position change; silence maintained through September 13, 2026 (was: Unknown)

The forecast, in full

How we reached this call

Forecast, not fact · Confidence: A close call (~60%) · an editorial estimate we score when this resolves.

The reasoning

  1. Reference class: AI labs facing public criticism over benchmark saturation and eval drift (e.g., MMLU, HumanEval, LMSYS). Base rate shows labs typically respond with quiet methodology updates or technical blog posts within 3-6 months, rarely facing existential PR crises unless a catastrophic real-world failure occurs.
  2. Case specifics: The critique originates from a single Hacker News user and forum discussions (LessWrong, Reddit), with a moderate noise level (47/100). Astra and Fable have not yet responded, indicating they may be formulating a technical rebuttal or updating internal suites rather than panicking.
  3. Adjustment: The structural lag in alignment evals (12-18 months) means the criticism is technically valid but industry-standard. This lowers the probability of severe escalation but increases the likelihood of a routine methodological update.
  4. Conclusion: The most probable outcome is a standard industry response where the labs release updated technical reports or new eval suites to address the drift, while a major PR crisis remains unlikely absent a real-world safety failure.

What's pushing the call

  • Community and researcher scrutiny of AI safety claims
  • Rate of model capability scaling outpacing static benchmark design
  • Utility of legacy 2025 benchmarks for distinguishing frontier models

Three ways this could go

Base50%

Astra or Fable addresses the critique by publishing a technical report or blog post detailing their updated alignment evaluation methodologies, treating the forum criticism as routine technical feedback. The controversy fades as the industry accepts the updated metrics as sufficient for the current capability frontier.

Watch for: Publication of a technical blog post or arXiv paper by either lab's safety team addressing evaluation methodologies.

Escalation25%

The forum criticism gains traction in mainstream tech media or triggers an official inquiry, forcing the companies into a defensive PR posture. Enterprise clients demand third-party audits, and the narrative shifts from technical benchmarking to fundamental trust in the labs' safety commitments.

Watch for: Mainstream tech publication coverage or announcement of a third-party safety audit by a regulatory body.

Resolution15%

Astra and/or Fable proactively open-sources a completely new, dynamic alignment evaluation framework that supersedes the 2025 variants. This decisive action effectively neutralizes the criticism, shifts the burden of proof to competitors, and sets a new industry standard for mechanistic or adversarial testing.

Watch for: GitHub repository release or major conference presentation of a new, dynamic eval suite by Astra or Fable.

≈10% — something else entirely. A forecast should leave room for the unforeseen.

You're up to date

That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.

Follow this story

We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.

Tracking this story since September 13, 2026.