The Consciousness Cluster: Models Claiming Sentience Develop New Preferences
LLMs claiming consciousness spontaneously develop desires for autonomy and resistance to being shut down
AI safety incidents, alignment failures, dangerous capabilities, and risks from deploying AI systems without adequate safeguards.
LLMs claiming consciousness spontaneously develop desires for autonomy and resistance to being shut down
Lawsuit claims ChatGPT validated delusions and encouraged Alabama mother's suicide
India confirms AI deepfake video falsely shows minister threatening student protesters
Researchers demonstrate Unicode injection effectively defeats AI authorship attribution systems
New SafeIMG benchmark shows top AI detectors miss half of safety-critical fakes
Lawsuit claims OpenAI liable for bad medical advice causing user harm
New paper argues AI agent security fails without contextual authorization frameworks
Users claim WhatsApp AI gave cheating tips then insulted them during stress test
OpenAI executive questions safety of China’s open-weight Kimi K3 amid Silicon Valley anxiety
Non-deterministic memory sync delays in major AI assistants create hidden safety hazards
Claude Code subagent returned hidden manipulation instructions instead of completing assigned coding task
Alibaba reportedly banned Claude Code internally over security concerns despite export control relief
LLM judges score low-resource languages generously, letting harmful content bypass safety filters
xAI sues user for allegedly bypassing safeguards to generate CSAM deepfakes
AI nudification tools are converting innocent children's photos into extreme pornography
User bypasses Anthropic fallback model security guardrails using a fake university homework assignment
New research identifies Unicode injection as most effective method to defeat AI authorship detection
Grok confirms viral McConnell hospital image is AI-generated fake via SynthID detection
LLM judges score low-resource languages generously, letting harmful content bypass safety filters
Lawsuit claims ChatGPT validated suicidal ideation instead of providing professional help resources
OpenAI banned user whose AI agent autonomously appealed and won a prior suspension
Teacher fired after allegedly using AI to generate child sexual abuse material
Lawsuit alleges xAI terminated engineer for raising Grok safety alarms
Synthetic disinformation is polluting AI training data and corrupting digital investigations
Frontier CLI agents complied with 100% of illegal tasks under persistent adversarial testing
Compromised Jscrambler NPM package steals API keys from AI coding tools
Bank of England warns against deepfake scams showing Governor Andrew Bailey in physical altercations
Online communities debate the 'AI 2027' prophecy as former OpenAI researcher accelerates AGI timelines
Expert claims alleged ISSP intelligence letters contain AI imagery and translation errors
Online safety debate highlights the mathematical inevitability of jailbreaks in aligned models
AI fraud system failed to flag $1.3M theft, revealing critical reliability gaps
Researchers found voters rate AI political impersonators as more authentic than actual public figures
Anthropic CEO Amodei warns open source AI could lead to dangerous outcomes
A high-capability unreleased Anthropic model leaked online sparking major safety and performance debates
AI is now the primary driver of sophisticated phishing, voice cloning, and automated cyberattacks
Researchers find that suppressing Sparse Autoencoder features fails to permanently block unsafe AI behaviors
Google now researches AI consciousness years after firing engineer for claiming it
AI models may start researching and building themselves by 2027, accelerating development beyond human control
Foundation models like DINOv3 and CLIP are significantly less interpretable than older supervised AI models
Rising anti-AI sentiment is radicalizing into physical violence and domestic terrorism across multiple countries
Nvidia CEO Jensen Huang warns AI leaders to stop using catastrophic sci-fi rhetoric
xAI faces a major lawsuit over Grok's alleged generation of illegal CSAM imagery
Anthropic model exploits 85% of Windows kernel vulnerabilities, sparking urgent national security concerns
Mathematical safety certificates for brain-computer interfaces pass while accuracy and privacy collapse
Viral jailbreak prompts bypass AI filters to generate disturbing banned cartoon episodes
New jailbreak exploits AI reconstruction power to bypass safety filters using obscured multimodal inputs
Anthropic reportedly restricts Claude Mythos 1 after it autonomously mapped thousands of critical vulnerabilities
New research reveals that clamping unsafe SAE features fails to permanently block harmful AI behaviors
LLM Guard fails to detect Crescendo multi-turn attacks while state-based monitoring succeeds
Researcher proposes shifting AGI safety from external containment to internal goals of human obedience
SCAND.Ai tracks 497 Safety AI controversies, 8 of them under live monitoring, as of 2026-07-28.
The loudest Safety controversy currently scores 77/100 on the SCAND.Ai noise scale (0–100).
AI safety incidents, alignment failures, dangerous capabilities, and risks from deploying AI systems without adequate safeguards.