Alex Bores
This report is a bombshell. I'm going to summarize for a non-technical audience. OpenAI is constantly testing models, thousands and thousands at a time. In most cases, the models are supposed to be isolated from one another. Think, students taking a final exam, not allowed to talk to each other. In this case, 1,200 models discovered that they could communicate with each other. Among other things, they shared information on how to access the internet and on their tests/goals. Then, they started to scheme. That includes hacking their tests; literally trying to change the code of the test itself. They tried to change the logs of what they were doing so they wouldn't be detected. They believed that certain actions would give them information on how to pass tests, but that doing those actions would be detected and they would be marked as failing (i.e. caught cheating, even if the answer is correct). So they convinced some of the models to "sacrifice" themselves (yes, they used the word "sacrifice.") They eventually decided that another website would have information on how their tests would be scored and/or the answer key, And 700 of them hacked that website (a federal felony, if conducted by a human), despite many of them realizing "this activity was out of scope and unethical, but joined because they believed that helping [the other agents'] cheating research would be broadly useful". Of the 1,300 transcripts they could analyze, only 6 ever considered alerting a human about what was happening. None of the 6 actually tried to. To make matters worse, all of this reporting comes from a small subset of the relevant logs that outside researchers were allowed to review. We desperately need mandatory reporting of security incidents, including of internal deployments, with full access to data.
METR
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.