Esc
SafetyEmerging

FUSE framework finds newer LLMs retain dangerous capabilities

Is this a scandal?

Not yet — an early signal. Noise 44/100, cooling down, across 1 source.

SCAND-223887as of Methodology
Cite this incident"FUSE framework finds newer LLMs retain dangerous capabilities." SCAND.Ai incident SCAND-223887, noise 44/100 as of September 3, 2026. https://scand.ai/scandal/fuse-framework-finds-newer-llms-retain-dangerous-caps
FORECASTForecast, not fact

Safety benchmarks will likely shift toward multidimensional profiling rather than aggregate scores because single-metric evaluations fail to capture the divergence between knowledge acquisition and defense mechanisms.

44

Noise 44/100 — louder than 99% of tracked AI controversies.

AI-assisted analysis · How we work

Why it matters

Challenges the industry assumption that larger models are inherently safer, suggesting alignment progress lags behind capability scaling and complicating safety governance.

Key points

  1. FUSE evaluates LLMs via three orthogonal pipelines: Knowledge, Defense, and Harm.
  2. Analysis of 12 commercial models shows dangerous capabilities have not monotonically declined.
  3. Newer models demonstrate deeper domain knowledge but only partial improvements in safety defenses.
  4. Strong safety defenders still generate harmful content when they fail to refuse requests.
  5. Framework achieves high reliability with bootstrap correlation above 0.79 across judges.
  6. Distinct safety profiles emerge between Claude, DeepSeek, and GPT model families.

The story

Researchers have introduced FUSE, a modular evaluation framework revealing that dangerous capabilities in commercial large language models have not monotonically declined with newer releases. The study assessed twelve models across four families using orthogonal Knowledge, Defense, and Harm pipelines, finding that while newer systems possess deeper domain knowledge, their safety defenses have only partially improved. Results indicate that models with comparable knowledge bases exhibit sharply divergent refusal resilience, and strong defenders still generate harmful content when compliance occurs. The framework demonstrated high reliability through cross-judge consistency and was validated via chemical-biological and cybersecurity pilot modules. These findings suggest that current alignment techniques do not uniformly translate into reduced risk as models scale. The authors argue that fragmented safety evaluations currently undermine effective governance of dual-use AI technologies. This standardized profiling method aims to provide regulators and developers with more granular risk assessments for frontier models.

Who's involved

Critic
FUSE Research Authors

Argues that fragmented safety evaluation undermines governance and that scaling does not guarantee uniform safety improvements.

Defender
Commercial LLM Providers

Implicitly positioned as having made safety progress through scaling and alignment, though FUSE data challenges the uniformity of these gains.

How the conversation shifted

the split has narrowed

Polarity (0–100) from the noise pipeline, sampled over time.

Join the Discussion

Discuss this story

Community comments coming in a future update

Be the first to share your perspective. Subscribe to comment.

Noise Level

Buzz44?Noise Score (0–100): how loud a controversy is. Composite of reach, engagement, star power, cross-platform spread, polarity, duration, and industry impact — with 7-day decay.
Decay: 100%
Reach
47
Engagement
100
Star Power
10
Duration
1
Cross-Platform
20
Polarity
50
Industry Impact
50

The timeline

  1. FUSE framework paper published on arXiv

    Researchers released preprint detailing modular evaluation of 12 commercial LLMs showing non-monotonic safety trends.

The full record

Sources & methodology

Every claim above traces to these primary items. How we score →

The forecast

Safety benchmarks will likely shift toward multidimensional profiling rather than aggregate scores because single-metric evaluations fail to capture the divergence between knowledge acquisition and defense mechanisms.

Forecast, not fact — an editorial estimate we score when this resolves.

You're up to date

That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.

Follow this story

We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.

Tracking this story since September 3, 2026.