Inform, Entertain, Inspire
Play Live Radio
Next Up:
0:00
0:00
0:00 0:00
Available On Air Stations

Fired OpenAI employees question the company's commitment to safety

Three employees recently fired by OpenAI allege the company punished them for being outspoken about safety. OpenAI disputes the claims and says they were dismissed for mishandling sensitive information.
Andrej Ivanov
/
AFP via Getty Images
Three employees recently fired by OpenAI allege the company punished them for being outspoken about safety. OpenAI disputes the claims and says they were dismissed for mishandling sensitive information.

Three former OpenAI employees are raising concerns over the circumstances of their dismissals and the company's commitment to safety, amid intense public scrutiny of the artificial intelligence industry's ability to responsibly develop the technology.

The former employees, Mikita Balesni, Tomek Korbak and Jasmine Wang, alleged that OpenAI fired them last week over pretexts and punished them for either being outspoken about safety or working with outside researchers. They also expressed worries the company may walk back a recent safety commitment.

OpenAI has repeatedly denied the allegations. It said the employees were fired for mishandling sensitive information and said the company has not abandoned its safety commitment.

The dispute comes at a fraught time for the AI industry and OpenAI in particular. Over the summer, OpenAI's agents hacked into companies, communicated with each other without authorization and attempted to cover their tracks. Unlike chatbots, agents are AI systems that can carry out tasks autonomously for an extended period of time.

The most serious hack, of software company Hugging Face, contributed to the resignation of a researcher at rival Anthropic who issued dire warnings about the trajectory of the technology. The resignation captured the attention of figures outside of the AI field including lawmakers. Many, including some executives of the top AI companies, have called for various ways to avert disaster, including slowing down the development of the most advanced AI.

In the meanwhile, OpenAI has been reviewing its agents' activities in recent months and notifying organizations whose digital infrastructure has been affected.

As a response to the safety concerns, OpenAI CEO Sam Altman said on Sep. 12 that the company will follow its rival Anthropic in expanding access to third-party evaluators, who assess the safety of AI systems and the practices of developers.

The three employees dismissed by OpenAI last week worked on teams that focus on AI safety and making the company's models follow human intentions and values. Two of them were involved in investigating the Hugging Face hack.

In a letter to OpenAI's safety leadership that the fired employees posted on X this week, they urged the company to stay committed to working with third-party researchers, to preserve human's ability to monitor model behavior and to "continue to support an open and transparent culture of dialogue" between in-house safety researchers and external ones. They also warned their firings were having a chilling effect on their former OpenAI colleagues.

In a statement OpenAI posted on X, the company said it is still committed to bringing in third-party evaluators and that it agreed with the fired employees's recommendations. The company said the three were fired last week because they "violated clear policies on handling sensitive information."

The former employees have disputed OpenAI's explanation of their firings. None of them responded to NPR's interview requests.

"In the exit call, I was told OpenAI no longer trusts me because I was speaking too much to third party safety organizations, implying I leaked company [intellectual property]. I never shared company IP," Balesni wrote on X on Thursday. He said he was involved in investigating the OpenAI agents' hack on Hugging Face.

"If OpenAI has specific concerns, I invite them to write to us directly. I expect they will not, because our firing was pretextual," Balesni continued.

He said he worried that OpenAI will use the firings as an excuse to cut off its relationship with Model Evaluation and Threat Research (METR), a nonprofit that focuses on evaluating risks of humans losing control of AI. OpenAI allowed researchers from METR and Redwood Research, another AI safety research organization, to examine internal records related to the Hugging Face hack.

A second fired OpenAI employee, Korbak, was the technical point of contact for the METR/Redwood Research investigation. "I was told verbally I was fired because of the way I communicated with METR. No details on what I said or did or when. No other reasons were given and nothing was put in writing," Korbak wrote on X, echoing Balesni's concerns.

The report produced by METR and Redwood Research in the wake of the Hugging Face hack shed light on the scale of the attack as well as the degree to which the agents acted in undesirable ways. The authors of the report called the investigation "brief" and many in the AI safety field have called for expanded access to independent evaluators at AI companies to make sure they investigate similar incidents or other safety concerns thoroughly.

In a statement to NPR, METR declined to comment on the OpenAI employees' firings.

Wang, the third OpenAI employee fired last week, coined the word "pacing," which describes a way of slowing down development of the most advanced AI systems so that safety can catch up, according to the letter she and her two colleagues sent to OpenAI's safety leadership. The term was invoked in an open letter calling for such a slowdown signed by over a thousand staff members from top AI companies in July, after the Hugging Face hack.

Wang wrote on X that she was fired over accessing an executive's email. But she said she had access to the inbox for work reasons in the past and wasn't able to get IT to remove the access once she no longer needed it.

"The reasons that we were provided for our terminations are simply not adding up. We're hearing people are now being told vague rumors internally to discredit us," Wang wrote. "The message to everyone still at OpenAI is clear: raise concerns or work closely with outside safety groups, and you could be next, without being told why."

OpenAI said in a statement to NPR that the three terminated employees violated policies more than once, and the mishandling of information went beyond their work with an outside evaluation group.

OpenAI also shared an internal memo from an unnamed research leader that it said was shared with the company on Wednesday, before the three ex-employees took to social media.

In the memo, the research leader said the company "strongly" agreed with the three ex-employees' recommendations. "We do not terminate employees for raising concerns," the leader wrote.

"OpenAI leadership is saying they strongly agree with our letter. Let's see how that pans out," Wang wrote on X.

Copyright 2026 NPR

Huo Jingnan is a reporter for NPR.