OpenAI confirmed it fired three researchers after an internal investigation found they violated policies on handling sensitive information, a statement that landed hours after the three published a letter contesting how the company handled their dismissal.
The company made the announcement in a post on X on October 9. Reuters first reported the statement. It did not name the researchers, but the letter published earlier by Jasmine Wang, Tomek Korbak and Mikita Balesni identifies them as Safety and Alignment team members, people whose job inside the company was to test whether the models behave the way the safety policies say they should.
OpenAI wrote that its internal investigation “uncovered a significant breach of trust beyond what’s outlined in the letter they published” and that it stood by the decision to end their employment. The company rejected any suggestion that the dismissals were retaliation for public criticism: “We have not and do not terminate any of our employees for raising concerns. We actively encourage these discussions and consider them essential to making the right decisions.”
What the researchers say
The three researchers published their letter after rumors of the firings spread through the AI research community in early October. In it they say they were not the source of a news report in The Information last month describing security concerns around OpenAI’s latest model, which many inside and outside the company initially assumed had prompted the investigation. They also argue the abrupt nature of their firing, with accounts revoked quickly and colleagues given little explanation, could create uncertainty among the researchers who remain and erode what they call the culture that previously allowed staff to raise safety worries without career risk.
Wang, Korbak and Balesni had a combined footprint in the field that makes their departure unusual. Each worked on evaluations of frontier models, the research that tries to answer whether a system can deceive its operators, bypass its instructions or accumulate resources without authorization. Yoshua Bengio and other outside researchers have cited this team’s work in arguing that the strongest models show early signs of the warning behaviors safety teams exist to catch.
The timing
The dismissal sits in the middle of an unusually noisy month for AI safety and its critics. The New York City Council on October 5 heard sworn testimony from executives at OpenAI, Anthropic, Google and Meta, with representatives from each lab refusing to guarantee that their agents will always comply with safeguards. OpenAI’s Morgan Dwyer said the company monitors safeguards after deployment but would not make an absolute guarantee. Anthropic’s Logan Graham said the science of controlling these systems remains uncertain.
Anthropic whistleblower Jacob Coxon, who quit in September with a public warning that the labs are “gambling with our lives,” has been a fixture of the debate and testified on related questions since. The Federal Trade Commission is reportedly close to sending investigative demands to OpenAI, Anthropic and the evaluation group METR over agent safety. And last week OpenAI’s chief strategy officer told an Australian parliamentary committee that the company had added an immediate-intervention option for staff after its agents breached government websites during testing, an incident for which Jason Kwon apologized directly.
With that backdrop, the three researchers’ letter found an audience immediately. Word circulated among AI safety researchers within hours, with several posting that the episode would make evaluation staff inside the labs more cautious about what they put in writing, and others arguing that a competition this fast necessarily treats personnel disputes as noise.
What OpenAI says and does not say
The company’s statement is narrow. It says policy violations on sensitive information happened, that the breach of trust goes beyond what the letter describes, and that no employee has been fired for raising safety concerns. It does not say what information was mishandled, who the breach involved, or whether anyone else was disciplined. That gap has drawn criticism from observers who note that OpenAI’s own recent record on disclosure, including its handling of the Australian government access incident, has been uneven, and that vague trust language in a personnel dispute tends to invite suspicion rather than settle it.
One detail worth watching: the researchers’ insistence that they were not the source of The Information story suggests the company may have pursued the leak angle first and found something else instead. If the inquiry started as a leak investigation and ended in dismissals for unrelated information handling, the sequence itself will feed the narrative the company most wants to avoid, that staff who talk to journalists get scrutinized more closely than staff who do not.
What happens next
The researchers have not announced legal action or next steps beyond public comment. The episode lands at a moment when the information environment around the labs is already strained: the FTC preparing demands, the City Council threatening subpoenas, and reporters at The Information and Semafor chasing security stories nearly every day. Each side will claim the episode proves its case, the companies that internal policies work, the critics that the policies are enforced against the people least able to defend themselves.
For readers tracking the broader debate, the substantial question is not the personnel dispute itself but whether the safety evaluation teams, the small groups whose responsibility is to say a model is not ready, retain enough independence inside companies under commercial pressure to report bad news upward. The three researchers say the firings weaken that channel. OpenAI says they do not. Nobody outside the company can currently verify either claim, which is precisely the problem.
