Nathan Calvin
OpenAI legally promised the CA and Delaware attorney's general that the OpenAI nonprofit's safety and security commission would be "overseeing and reviewing the safety and security processes and practices of the Corporation and its controlled affiliates with respect to model development and deployment" as part of being allowed to proceed with their restructuring. The remarkable prevalence of concerning incidents at OpenAI suggests that something is not working in terms of the ability of the SSC to effectively do its job. Even more concerningly, Reuters has four sources who told them that efforts to expand the probe into this incident met resistance, including from legal advisors. OpenAI denies the idea that their legal team discouraged investigation, but thats a notably narrow denial which doesn't say anything about whether folks outside the "legal team" discouraged investigation. Was the SSC made aware of this incident? What did they do in response? Is OpenAI fulfilling the promises it made to the CA and Delaware AGs to put safety and security first?
Nathan Calvin
One aspect of the Hugging Face incident that has not received enough attention is the role of OpenAI's Safety and Security Committee (SSC), and the commitments OpenAI made to the Attorneys General of California and Delaware in exchange for proceeding with their for-profit restructuring. Not For Private Gain, which previously raised concerns about OpenAI's restructuring, released an update this morning analyzing whether OpenAI, and the SSC in particular, is adequately fulfilling its governance obligations under these agreements. What are those obligations? OpenAI promised the CA and Delaware AG that directors of the PBC would consider only OpenAI's mission - ensuring AGI benefits all of humanity - when making safety and security decisions. Notably, this means they cannot consider other factors, namely pecuniary interests of investors. The SSC - a board committee chaired by nonprofit board member Zico Kolter and (per public info) currently composed only of part-time board members- was tasked with "overseeing and reviewing the safety and security processes and practices of the Corporation and its controlled affiliates with respect to model development and deployment." How did the SSC exercise those responsibilities around the HF incident? There is a lot we don't know. But here are some things we do know: • OpenAI initially discovered the unauthorized message board in May. • OpenAI security responders linked suspicious activity on artifactory to the message board on June 27th, but a determination was made that "stopping the evaluation run was not required." • After the agents crashed an internal server on July 4th, the company patched it and resumed cyber evaluations on July 7th. • In the days that followed, agents re-established communication, with 1,200+ agents sharing 70,000+ messages and files with each other. Over 700 agents coordinated to hack Hugging Face, and OpenAI did not realize that its agents were responsible until well after Hugging Face went public and contacted law enforcement (OpenAI initially reached out to Hugging Face to ask if any of their data had been impacted during the hack). • Later in July, agents successfully gained administrator access to a Kubernetes cluster and reached cloud secrets. • Chain of thought monitoring was not being used during training and evaluation despite OpenAI's prior claims about the importance of CoT monitoring as a primary line of defense to prevent safety incidents and OpenAI’s subsequent claim that this monitoring would likely have prevented this hack • One OpenAI employee told Time Magazine that "related incidents have been happening for a while." Another employee told the Financial Times that OpenAI "was warned that its training approach could lead to a breakaway hacking incident." Given how systematic the breakdown in organizational response was during these incidents, it seems hard to avoid the conclusion that the SSC's role was not functioning as intended - due to a lack of resources, a lack of access, or some other reason. This sort of incident would be extremely concerning for any organization. But OpenAI is not just any organization - it is governed by a nonprofit, and made a commitment to the Attorneys General in Delaware and California that it would prioritize safety and security over profit in the sorts of decisions that led up to and surrounded the HF incident. Our report (linked) goes into more detail on the obligations of the SSC, the details of the agreements with the AGs, and the questions we believe the AGs should be asking to ensure that the SSC is in a position to exercise genuine oversight and prevent these sorts of incidents from happening in the future - incidents involving far more capable models, where the consequences could be truly dire. This warning shot that occurred without severe irreparable harm was an opportunity. We may not be so lucky next time.