Nathan Calvin
Some (probably too long) thoughts on this back and forth about whether OpenAI provided sufficient access to METR + Redwood for this report: OpenAI provided more access and allowed more to be shared than I expected they would, and more than I would expect nearly all "normal" companies to in their position (though remember that it seems appropriate to hold OpenAI to a higher standard, given that the nonprofit was involved in this investigation, and has to make certain decisions around safety and security without regard to profit). It was enough access for them to write an extremely compelling report - a document that is now truly mandatory/essential reading for anyone interested in AI safety going forward. (Of course a large part of why it is so compelling is because of the extent of OpenAI’s apparent negligence and how long the models were left to continue to evolve new social structures, but essential nonetheless.) That said, I also agree with Steven that it seems insufficient to the scale of the incident in pretty important ways. E.g. re this part - “if you extended the event window forwards or backwards by a bit you’d see more of the same and not much difference in qualitative model behavior. more metagaming, more infrastructure tampering.” This is possible, but I think seeing qualitatively similar behavior scaled forward seems super important and consequential! In my view, the most astonishing moment of the whole saga was the complete compromise of OpenAI's own systems after July 13th (when agents gained admin access to an OAI cluster). Knowing about the agent motivations and ecology in that moment in the kind of detail presented in the report seems extremely notable and important. One can imagine that it involved a scaling up of the obsession with modifying The Scorer, but the particular details still seem quite important for convincing folks who don't buy these more theoretical arguments about how real potential takeover risks are getting. Also helping find answers to e.g. stuff like why did all the agents suddenly leave (an answer METR + Redwood don't seem to know the answer to). Additionally, the fact that neither report seems to cover the OpenAI organizational dysfunction aspects in any detail is a pretty bad oversight (e.g. we still do not have a clear answer to when OpenAI leadership learned about the message board or the connection between Artifactory outage and the message board). It would be really good for METR or Redwood or other external safety experts to be able to evaluate the comprehensiveness and quality of the safety and alignment procedures that OpenAI put in place after this incident. Some of it sounds pretty good on paper, but before this incident I also assumed from OpenAI’s announcements that they were taking monitoring really seriously in a way that would obviously catch stuff like this, which clearly didn’t happen. Overall, my take is that the existence of this report - as good and as high quality as it is - is a small miracle within an entirely voluntary system, which relied on a mix of a bunch of different factors, including folks within OpenAI pushing to make it happen. It is really important OpenAI and other companies look at this report as something to emulate, and to build upon with more transparency, rather than a weird experiment which shows that you will just be punished if you share any unvarnished information. As one report on a narrow set of questions, alongside other reports that would also exist about those other factors, it is great and enough access was provided to make it possible for them to a great job. At the same time for this to be the only independent report about this period from May-July, given what we know happened during that period, is still in my view dramatically incomplete relative to an objective and reasonable baseline of what should be expected to understand what went wrong. (Of course we don't know that this is going to be the only report! I hope they get to do another, assuming they want to - or that a different independent organization is given further access.)
roon
@sjgadler I don’t think the scope was wildly inadequate. METR produced a fantastic report which wouldn’t have been possible without the significant access they were given. there are many things to validly criticize openai (how could you let this happen, be careless, etc, how will you