Jasmine Wang, Tomek Korbak, and Mikita Balesni, the three safety researchers that OpenAI fired last week, have published an open letter denying the firm’s claims that they mishandled sensitive information outside of established company procedures and warned that their dismissal signals a chilling effect that will have ripple effects across the company’s culture.

“We have become concerned that internal and external communications around our firing have made our former colleagues afraid to speak and operate in ways that, until last week, were an integral part of working at OpenAI,” the researchers wrote Thursday in an open letter to OpenAI’s Safety and Security Committee, Safety Advisory Group, and Mission Advisory Council. 

The researchers were dismissed last week after allegedly sharing confidential company information with a third-party AI safety organization. OpenAI said at the time that they violated the company’s policies by “accessing and handling sensitive company information.”

“AI is not a normal technology, and OpenAI is not a normal company,” Wang, Korbak, and Balesni wrote. “Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them. The freedom to do so without fear, and to have well-defined internal procedures that enable this work, is itself an essential safety mechanism.”

They said that their firing represents a broader shift in the culture of OpenAI, one that used to encourage workers to “raise safety concerns and disagree openly.” They said employees are now “unclear on where they stand” when behavior that was allegedly normal a month ago is now suddenly grounds for dismissal. 

“Given the significant safety concerns surrounding the development of AI, employees must not be left working in an environment where fear and unclear rules stymie AI safety work and weaken third-party accountability,” they wrote. “Terminations such as ours, executed and communicated so abruptly, are chilling the open culture OpenAI has prized in the past.”

In the letter, the three denied involvement in a leak to The Information about less monitorable architectures in OpenAI’s newest models that make chain-of-thought reasoning more difficult to monitor. They also denied engaging with external parties outside the mandates of their jobs. 

OpenAI has not formally responded to the open letter, but shared with TechCrunch an internal memo attributed to a research leader, praising the three researchers’ contributions to AI safety and denying that they were fired in retaliation.

“I want to be very clear that these decisions were not about raising safety concerns or speaking out,” the memo reads. “We have always encouraged that and always will. We do not terminate employees for raising concerns.”

OpenAI did not directly address TechCrunch’s questions about which policies the researchers allegedly violated, the circumstances of their dismissal, or how the company protects employees who raise safety concerns and collaborate with external evaluators.

The firings have fueled speculation about their circumstances, particularly as OpenAI faces scrutiny over recent safety incidents involving rogue agents and leaks about its models.

The letter also addresses the researchers’ response to the Hugging Face incident, in which a swarm of agents broke out of their sandbox and breached external systems. The letter says that the incident and investigation was “without precedent,” meaning “internal policies were being developed in real time.” Due to the sensitive nature of the investigation, Korbak believed he was acting within OpenAI’s policies and norms by communicating closely with outside safety evaluators to build trust, per the letter. 

At the same time, Balesni was also working internally to address the growing AI monitorability problem, an effort the researchers say in their letter “can only succeed through extensive communication with external parties.” According to the letter, Balesni coordinated with and was supported by OpenAI board members and executives throughout his work.

“Throughout, Mikita checked in with his reporting line and took care to remove sensitive details from materials before sharing them,” the letter reads. “He acted throughout in good faith and within the company’s norms as they stood at the time.”

In a separate thread on X, Wang explained more details about her own dismissal, explaining that OpenAI told her she’d been fired because she accessed an executive’s email.

“OpenAI delegated that access to me for recruiting,” she wrote. “When I no longer needed it, I asked IT to remove it. They did not action my request, I couldn’t remove it myself, and the inbox was combined in an indistinguishable way in my phone’s mail app. When I opened a sensitive email by mistake, I told the executive within minutes and asked IT again. None of this was hidden.”

Wang went on to say that the reasons behind the terminations are “not adding up,” and that she and her colleagues are “not the first to be pushed out of OpenAI under suspicious circumstances.”

The researchers called on OpenAI to adhere to its public commitments to embed third-party safety auditors within the organization, to preserve monitorability of frontier models, and “continue to support an open and transparent culture of dialogue between safety researchers and the rest of the safety ecosystem.”

OpenAI agrees with their recommendations, per the memo.

“Unless the employees take a stand now against this kind of maneuver, I am concerned we will not be the last,” Wang said. “The message to everyone still at OpenAI is clear: raise concerns or work closely with outside safety groups, and you could be next, without being told why. You can’t build AGI safely if the people closest to the risks are afraid to speak.”

Topics

, ,