Nathan Calvin

Nathan Calvin

@_nathancalvin · Twitter ·

There are three aspects of the HF OAI incident: 1. Agent behavior/alignment 2. Security/monitoring/controls 3. Organizational culture/negligence METR did a great report on #1. Zack and others are saying that 2 and 3 are also important and could have prevented the incident even if the agents were misaligned (which seems certainly true at least for this incident, even if its unclear to me whether it will be true for all future incidents). The solution to me clearly feels like this should be a yes AND situation - where you can give criticism on OpenAI for not releasing as much on two and three or allowing enough access to other independent experts on two and three well, while acknowledging that the insights we learned from one are genuinely fascinating and important. There are multiple things you can say that are true about this incident. You can say that OpenAI displayed remarkably basic lapses in monitoring and security, and that their organizational culture seemed to not be remotely effective for surfacing and addressing what should have been blaring red flags. And that if they had done so, regardless of all this newfangled fancy AI stuff, Hugging Face would not have gotten hacked and their cluster would not have gotten taken down. This is all true, to the best of my knowledge. At the same time, it is also true that OpenAI's apparent negligence allowed perhaps the most high fidelity and capable "model organisms" for misalignment the field has ever seen. This is extremely important to study, and the behaviors the agents engaged in are objectively fascinating and insane! I don't think anyone who has actually read the METR report and spent time looking at the raw transcripts could agree (and tbh I think some of the complaints may have come from folks who didn't read the report). Refusing to look at the chain of thought or engage with the apparent drives or goals of the agents, and only focusing on the prosaic cybersecurity issues seems like a massive mistake to me. Anyway - to me the answer seems very clear! This incident was about prosaic cybersecurity, the incentives and decision making of AI executives, AND misalignment and loss of control of remarkably capable AI systems. We can say it is about all three! Saying that it is about prosaic cybersecurity and incentives and therefore not about misalignment is silly! It's ok to criticize OpenAI for not talking enough about prosaic cybersecurity and incentives because those don't make their models look super capable and wild, and instead just makes them look incompetent, but its also true that their models are super capable and wild and it is appropriate to find this disconcerting! It's also absolutely not the case that OpenAI is going to be the only company with sloppy prosaic cybersecurity and bad incentives, so seeing what happens when you combine increasingly advanced AI systems with those things seems very notable!

Zack Korman

Zack Korman

The independent review of the OpenAI Hugging Face incident, supposedly a watershed moment in cybersecurity, wasn't done by a cybersecurity firm and the authors have no cybersecurity experience. That's bad. Here's my new video.