FUSE framework finds newer LLMs retain dangerous capabilities
Is this a scandal?
Not yet — an early signal. Noise 44/100, cooling down, across 1 source.
Safety benchmarks will likely shift toward multidimensional profiling rather than aggregate scores because single-metric evaluations fail to capture the divergence between knowledge acquisition and defense mechanisms.
Noise 44/100 — louder than 99% of tracked AI controversies.
Why it matters
Challenges the industry assumption that larger models are inherently safer, suggesting alignment progress lags behind capability scaling and complicating safety governance.
Key points
- FUSE evaluates LLMs via three orthogonal pipelines: Knowledge, Defense, and Harm.
- Analysis of 12 commercial models shows dangerous capabilities have not monotonically declined.
- Newer models demonstrate deeper domain knowledge but only partial improvements in safety defenses.
- Strong safety defenders still generate harmful content when they fail to refuse requests.
- Framework achieves high reliability with bootstrap correlation above 0.79 across judges.
- Distinct safety profiles emerge between Claude, DeepSeek, and GPT model families.
The story
Researchers have introduced FUSE, a modular evaluation framework revealing that dangerous capabilities in commercial large language models have not monotonically declined with newer releases. The study assessed twelve models across four families using orthogonal Knowledge, Defense, and Harm pipelines, finding that while newer systems possess deeper domain knowledge, their safety defenses have only partially improved. Results indicate that models with comparable knowledge bases exhibit sharply divergent refusal resilience, and strong defenders still generate harmful content when compliance occurs. The framework demonstrated high reliability through cross-judge consistency and was validated via chemical-biological and cybersecurity pilot modules. These findings suggest that current alignment techniques do not uniformly translate into reduced risk as models scale. The authors argue that fragmented safety evaluations currently undermine effective governance of dual-use AI technologies. This standardized profiling method aims to provide regulators and developers with more granular risk assessments for frontier models.
Who's involved
Argues that fragmented safety evaluation undermines governance and that scaling does not guarantee uniform safety improvements.
Implicitly positioned as having made safety progress through scaling and alignment, though FUSE data challenges the uniformity of these gains.
How the conversation shifted
Polarity (0–100) from the noise pipeline, sampled over time.
Noise Level
The timeline
FUSE framework paper published on arXiv
Researchers released preprint detailing modular evaluation of 12 commercial LLMs showing non-monotonic safety trends.
The full record
Sources & methodology
Every claim above traces to these primary items. How we score →
The forecast
Safety benchmarks will likely shift toward multidimensional profiling rather than aggregate scores because single-metric evaluations fail to capture the divergence between knowledge acquisition and defense mechanisms.
Forecast, not fact — an editorial estimate we score when this resolves.
That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.
Follow this story
We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.
Tracking this story since September 3, 2026.
Join the Discussion
Discuss this story
Community comments coming in a future update
Be the first to share your perspective. Subscribe to comment.