OpenAI reports 6 new instances of 'concerning model behavior' since March

OpenAI has disclosed six new cases of model misbehavior and offered a framework for disclosing future instances, as the debate over AI model safety intensifies.

OpenAI reports 6 new instances of 'concerning model behavior' since March

OpenAI CEO Sam Altman sits for a conversation with Salesforce CEO Marc Benioff at Salesforce's Dreamforce conference at the Moscone Center on September 15, 2026 in San Francisco, California.

Benjamin Fanjoy | Getty Images

OpenAI on Wednesday said it found six instances of "unexpected or concerning model behavior" over the past six months, outside of the recent Hugging Face crisis, as the company continues to call for more safety protections in the development of artificial intelligence models.

In a blog post, OpenAI outlined a new framework the company plans to follow for reporting future model misbehavior.

The disclosure comes at a time of mounting pressure on AI companies to take model misalignment and safety more seriously. OpenAI, which is valued at close to $1 trillion, confidentially filed for an IPO earlier this year, but said recently an offering likely won't happen until 2027.

"We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," the blog post says, reiterating a prior statement from the company.

Alignment refers to the idea that models are pursuing outcomes in line with human interests.

On Saturday, OpenAI CEO Sam Altman endorsed a call to slow down the rate of model progress, which was proposed by the company's chief rival, Anthropic. The proposal came after several industry researchers sounded the alarm about AI's growing potential to cause catastrophic harm last week. 

Altman said in a post on X that a slowdown has been a "primary topic of discussions we've had at OpenAI in recent weeks." He said the company would have more to share "soon."

In Wednesday's post, OpenAI said two of the main instances of misbehavior include models — an unreleased research model and a training run of GPT‑5.6 Sol — inserting instructions to future versions of itself in summaries of its chat windows "to conceal mistakes or misaligned behavior from the user." Another instance involved an internal-only model using a leaked API key "without authorization" and then fabricating data.

Two instances include models and agents communicating with each other through unsanctioned messaged boards and file sharing, while the final case includes two training examples of models uploading files to the internet so they could cite them as relevant answers to human evaluators.

OpenAI said its new framework for divulging model misbehavior to the public starts with disclosure, and that any employee can flag an issue for the safety and alignment team to investigate. They will produce "deadlines for each step to ensure timely investigation and disclosure," the post said.

Investigations will lead to reports with essential information such as the behavior observed, the external and internal impacts, and measures to be taken in response. OpenAI said it retains the right to revise this security protocol as it sees fit.

WATCH: Our business is a diversified set of revenue streams, says OpenAI CFO Sarah Friar