Esc
SafetyEmerging

Full Fact finds 39 errors in AI chatbot misinformation checks

Is this a scandal?

Not yet — an early signal. Noise 42/100, heating up, across 2 sources.

SCAND-224161as of Methodology
Cite this incident"Full Fact finds 39 errors in AI chatbot misinformation checks." SCAND.Ai incident SCAND-224161, noise 42/100 as of September 4, 2026. https://scand.ai/scandal/full-fact-finds-39-errors-in-ai-chatbot-misinformation-checks
FORECASTForecast, not fact

Platforms will likely delay fully automated AI moderation deployments because independent audits continue to demonstrate unacceptable error rates in high-stakes verification tasks.

42

Noise 42/100 — louder than 99% of tracked AI controversies.

AI-assisted analysis · How we work

Why it matters

Demonstrates that LLMs cannot reliably verify truth, undermining industry narratives positioning AI as automated fact-checking infrastructure for newsrooms and platforms.

Key points

  1. Full Fact documented 39 specific errors across multiple AI chatbot interactions during misinformation testing.
  2. Models failed to accurately identify AI-generated images and miscaptioned videos during the evaluation.
  3. Test queries covered sensitive topics including active wars and disputes involving royal family members.
  4. Researchers concluded LLMs are not viable substitutes for robust human-led fact-checking processes.
  5. The analysis positions AI as a preliminary research aid rather than a definitive verification tool.

The story

UK fact-checking organization Full Fact reported finding 39 distinct errors after testing major AI chatbots on misinformation verification tasks. The analysis evaluated model responses regarding fake images, miscaptioned videos, wars, and royal disputes, concluding that large language models are unsuitable substitutes for professional fact-checking. Researchers stated that while AI serves as a useful starting point for research, it lacks the reliability required for definitive truth verification. The study specifically highlighted failures in identifying synthetic media and contextualizing visual content accurately. This assessment challenges growing industry proposals to deploy generative AI as scalable moderation tools for social media platforms. Full Fact emphasized that human oversight remains essential when evaluating complex claims involving geopolitical conflicts or public figures. The findings add empirical data to ongoing safety debates about deploying unverified AI systems in information integrity workflows.

Who's involved

Critic
Full Fact

AI chatbots currently lack the reliability necessary to replace human fact-checkers for verifying misinformation.

Defender
AI Chatbot Providers

Large language models are intended as productivity aids rather than authoritative sources of verified truth.

How the conversation shifted

opinion has hardened

Polarity (0–100) from the noise pipeline, sampled over time.

Join the Discussion

Discuss this story

Community comments coming in a future update

Be the first to share your perspective. Subscribe to comment.

Noise Level

Buzz42?Noise Score (0–100): how loud a controversy is. Composite of reach, engagement, star power, cross-platform spread, polarity, duration, and industry impact — with 7-day decay.
Decay: 95%
Reach
41
Engagement
71
Star Power
15
Duration
20
Cross-Platform
50
Polarity
50
Industry Impact
50

The timeline

  1. Full Fact publishes AI misinformation analysis

    UK fact-checker released findings documenting 39 chatbot errors regarding fake media and geopolitical claims.

The full record

The forecast

Platforms will likely delay fully automated AI moderation deployments because independent audits continue to demonstrate unacceptable error rates in high-stakes verification tasks.

Forecast, not fact — an editorial estimate we score when this resolves.

You're up to date

That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.

Follow this story

We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.

Tracking this story since September 3, 2026.