Esc
EthicsCase Closed

Researchers introduce MiJaBench exposing demographic safety gaps in LLMs

Is this a scandal?

No longer — the story has resolved. Noise 3/100, cooling down, across 0 sources.

SCAND-161892as of Methodology
Cite this incident"Researchers introduce MiJaBench exposing demographic safety gaps in LLMs." SCAND.Ai incident SCAND-161892, noise 3/100 as of September 11, 2026. https://scand.ai/scandal/researchers-expose-selective-llm-safety-gaps-with-mijabench
FORECASTForecast, not fact

AI developers and evaluators will likely transition away from aggregate safety metrics toward more granular, demographic-specific benchmarks like MiJaBench. This will push frontier model providers to adopt generalized alignment techniques, such as DPO, to ensure uniform safety guardrails before deploying public API updates.

3

Noise 3/100 — louder than 95% of tracked AI controversies.

AI-assisted analysis · How we work

Why it matters

The discovery of a 'Selective Safety Trap' challenges the reliability of current LLM alignment, showing that safety protocols protect some groups while leaving others vulnerable to identical attacks.

Key points

  1. The new MiJaBench framework contains 43,961 controlled jailbreaking prompts across 16 minority groups in English and Portuguese.
  2. Safety guardrail defense rates varied by up to 42% within individual LLMs depending entirely on the demographic group targeted.
  3. The safety disparities persist across model sizes, architectures, and languages, and are actually amplified by model scaling.
  4. Applying targeted Direct Preference Optimization (DPO) allowed a 1B-parameter baseline model to achieve zero-shot safety generalization to unseen demographics.

The story

Researchers have introduced MiJaBench, a bilingual adversarial benchmark designed to systematically evaluate safety disparities across different demographic groups in Large Language Models (LLMs). The study, which analyzed 14 state-of-the-art models, revealed a systemic 'Selective Safety Trap' where models successfully block harmful prompts targeting certain demographics but remain highly vulnerable to identical attacks targeting others. Defense rates within the same model fluctuated by up to 42% depending solely on the targeted minority group, a disparity that persisted across various model architectures and languages. The authors demonstrated that current alignment techniques tend to learn group-specific safeguards rather than generalized harm prevention. To address these vulnerabilities, researchers proposed targeted direct preference optimization (DPO) on a baseline model, achieving zero-shot safety generalizations to unseen demographics and complex jailbreak strategies.

Who's involved

Critic
MiJaBench Research Team

Argues that current LLM safety evaluations create a dangerous illusion of universal protection by masking deep demographic inequalities in guardrail enforcement.

Neutral
LLM Developers

Providers of the 14 evaluated state-of-the-art LLMs whose models exhibited safety disparities under the selective safety trap testing.

How the conversation shifted

the split has narrowed

Polarity (0–100) from the noise pipeline, sampled over time.

Join the Discussion

Discuss this story

Community comments coming in a future update

Be the first to share your perspective. Subscribe to comment.

Noise Level

Quiet3?Noise Score (0–100): how loud a controversy is. Composite of reach, engagement, star power, cross-platform spread, polarity, duration, and industry impact — with 7-day decay.
Decay: 27%
Reach
0
Engagement
0
Star Power
10
Duration
0
Cross-Platform
0
Polarity
50
Industry Impact
50

The timeline

  1. Selective Safety Trap paper published

    Researchers release MiJaBench, exposing that LLM safety defenses vary up to 42% depending on the demographic group target.

The full record

What's being under-reported

No defender-side coverage yet

The critic side is sourced here; no defending voice has been captured yet.

  • Coverage: 0 social posts, 0 news-outlet items.
  • Voices: 1 critic, 0 defenders.

The forecast

AI developers and evaluators will likely transition away from aggregate safety metrics toward more granular, demographic-specific benchmarks like MiJaBench. This will push frontier model providers to adopt generalized alignment techniques, such as DPO, to ensure uniform safety guardrails before deploying public API updates.

Forecast, not fact — an editorial estimate we score when this resolves.

You're up to date

That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.