Esc

OpenAI denies using private math data in model training

Is this a scandal?

Not yet — an early signal. Noise 36/100, holding steady, across 1 source.

SCAND-231836as of Methodology
Cite this incident"OpenAI denies using private math data in model training." SCAND.Ai incident SCAND-231836, noise 36/100 as of September 12, 2026. https://scand.ai/scandal/openai-denies-using-private-math-data-in-model-training
FORECASTForecast, not fact

Expect updated terms of service explicitly defining de-identified data usage rights because ambiguity fuels ongoing IP litigation and erodes enterprise trust in AI platforms.

36

Noise 36/100 — louder than 99% of tracked AI controversies.

AI-assisted analysis · How we work

Why it matters

Clarifies boundaries between public outputs and private user data in AI training, setting precedent for intellectual property disputes involving proprietary research.

Key points

  1. OpenAI denies accessing Alpöge and Buckmaster's private data to solve the mathematical problem.
  2. Company admits de-identified usage data from researchers may have improved models generally.
  3. AI-generated proofs differ significantly from researchers' work on forced vs unforced Euler case.
  4. Statement distinguishes between direct data access for inference and aggregated training data.
  5. Response addresses IP concerns about AI leveraging non-public user inputs for capabilities.

The story

OpenAI stated it did not access private user data from mathematicians Levent Alpöge and Tristan Buckmaster to solve a mathematical problem recently solved by its agents. The company acknowledged that while no specific user data was viewed during the solution process, de-identified usage data derived from their product interaction may have contributed to general model improvements. OpenAI emphasized that its generated proofs differ significantly from the researchers' work, particularly regarding forced versus unforced Euler equations. This response addresses concerns about whether AI systems leverage non-public user inputs for capability advancement. The statement aims to distinguish between direct data access for inference and aggregated, anonymized data used in training pipelines. Researchers had questioned potential unauthorized use of their unpublished work after observing similar AI-generated results.

Who's involved

Critic
Levent Alpöge and Tristan Buckmaster

Questioned whether AI used their unpublished mathematical research without authorization

Defender
OpenAI

Denies accessing private user data for inference while acknowledging de-identified data may aid training

Join the Discussion

Discuss this story

Community comments coming in a future update

Be the first to share your perspective. Subscribe to comment.

Noise Level

Murmur36?Noise Score (0–100): how loud a controversy is. Composite of reach, engagement, star power, cross-platform spread, polarity, duration, and industry impact — with 7-day decay.
Decay: 81%
Reach
49
Engagement
43
Star Power
35
Duration
70
Cross-Platform
20
Polarity
50
Industry Impact
50

The timeline

  1. OpenAI issues public denial on Twitter

    Company responds to allegations of using private math research data, distinguishing inference access from training data usage

The full record

Sources & methodology

Every claim above traces to these primary items. How we score →

The forecast

Expect updated terms of service explicitly defining de-identified data usage rights because ambiguity fuels ongoing IP litigation and erodes enterprise trust in AI platforms.

Forecast, not fact — an editorial estimate we score when this resolves.

You're up to date

That's the complete picture as of — nothing more to know right now. We'll update this page the moment it changes.

Follow this story

We keep this page current — no need to check back. We'll send the next real change to your inbox, nothing else.

Tracking this story since September 8, 2026.