Esc
EthicsCase Closed

Efficiency Controversy Hits GLM 5.1 Over Excessive Chain-of-Thought Tokens

Is this a scandal?

No longer — the story has resolved. Noise 1/100, cooling down, across 0 sources.

SCAND-66871as of Methodology
Cite this incident"Efficiency Controversy Hits GLM 5.1 Over Excessive Chain-of-Thought Tokens." SCAND.Ai incident SCAND-66871, noise 1/100 as of July 25, 2026. https://scand.ai/scandal/glm-5-1-excessive-thinking-controversy
FORECASTForecast, not fact

Model developers will likely introduce 'thinking' limits or more aggressive pruning of reasoning paths to manage costs. In the near term, expect new benchmarks to emerge that measure 'intelligence per token' to penalize models that use brute-force reasoning.

1

Noise 1/100 — louder than 88% of tracked AI controversies.

AI-assisted analysis · How we work

Why it matters

The shift toward 'thinking' models raises questions about the economic viability and efficiency of open-weights models compared to proprietary giants. It highlights a potential trend where model performance is boosted through sheer token volume rather than architectural intelligence.

Key points

  1. GLM 5.1 is reportedly using excessive Chain-of-Thought tokens, reaching over 150,000 tokens for simple coding requests.
  2. Users report significant latency issues, with the model 'thinking' for 20 to 30 minutes before providing a final answer.
  3. Despite the high token overhead, the output quality remains inconsistent, with reports of basic programming errors like accessing protected members.
  4. The controversy highlights a shift in AI benchmarking where 'intelligence' may be tied to token volume rather than architectural efficiency.

The story

A controversy has emerged regarding the efficiency and cost-effectiveness of the recently released GLM 5.1 open-weights model. Users are reporting that the model's 'Chain-of-Thought' (CoT) reasoning process consumes an excessive number of tokens for relatively simple tasks, such as basic UI programming. Reports indicate that the model can spend upwards of 30 minutes and 150,000 tokens on a single prompt, frequently oscillating through internal 'corrections' before producing a final output. While the model is praised for its accessibility as an open-weights alternative to Claude and ChatGPT, critics argue that the high token consumption negates its price advantage. Preliminary user tests also suggest that despite the exhaustive 'thinking' phase, the resulting code still contains basic syntax errors and requires human intervention, sparking a debate on whether state-of-the-art performance is being artificiality inflated via brute-force token generation.

Who's involved

Critic
FPham (Reddit User/Developer)

Argues that the excessive token usage makes the model 'price-unsmart' and inefficient compared to Claude or ChatGPT.

Defender
Zhipu AI (GLM Developers)

Developing state-of-the-art open-weights models that utilize extensive reasoning to match proprietary performance.

Neutral
Open Source AI Community

Divided between appreciating the power of open-weights models and concerns over the hardware/cost requirements of running them.

Join the Discussion

Discuss this story

Community comments coming in a future update

Be the first to share your perspective. Subscribe to comment.

Noise Level

Quiet1?Noise Score (0–100): how loud a controversy is. Composite of reach, engagement, star power, cross-platform spread, polarity, duration, and industry impact — with 7-day decay.
Decay: 5%
Reach
0
Engagement
0
Star Power
15
Duration
0
Cross-Platform
0
Polarity
65
Industry Impact
75

The timeline

  1. Model Quality Self-Correction

    The model's internal reasoning admits it is 'overcomplicating' the task, leading to user skepticism about the SOTA claims.

  2. Token Usage Analysis

    Reports surface that the model consumed over 100k tokens before producing functional code.

  3. GLM 5.1 User Report Gains Traction

    A developer documents a 30-minute 'thinking' loop for a simple C++ UI component request.

The forecast

Model developers will likely introduce 'thinking' limits or more aggressive pruning of reasoning paths to manage costs. In the near term, expect new benchmarks to emerge that measure 'intelligence per token' to penalize models that use brute-force reasoning.

Forecast, not fact — an editorial estimate we score when this resolves.

You're up to date

That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.