Joshua Saxe
Just because AI safety has become incoherent as a research program -- (citing my own QT post below) -- doesn't mean we should become libertarian accelerationists. Quite the opposite; there are all sorts of present and future harms that couldn't be more important to mitigate. But there's a new framing that needs to happen that would break the inertia of the past seven years... it would dial down the emphasis on jailbreaks & guardrails (given the reality of uncoordinated diffusion of models that won't have jailbreaks & guardrails), dial up an emphasis on resilience given the fact of diffusion, and it would discount (but not totally ignore) epistemes based in thought experiments. It would see risks as interactions between models and society not simply reflective of in-lab demonstrations of deception and scheming (although these would be seen as among the moving parts of these more completely conceived objects of risk). It would be more interdisciplinary; instead of making assumptions (e.g. GDPval-maxxing == job loss) it'd upweight economists; instead of making assumptions like maxxing CyberGym = more cybercrime, it'd upweight cybersecurity practitioners. Funding sources and discursive centers of gravity would shift. With respect to the brilliant researchers there, a new safety paper from Anthropic would not be the highest degree centrality intervention in the safety discussion in any particular month. Instead, independent organizations with fewer conflicts of interests would take center stage; federal / government funding agencies and public research organizations (national labs and universities) would resume a more central place. This all sounds like a fantasy, but I think it's actually inevitable; real harms will be a rising tide that forces it to be so.
Joshua Saxe
AI safety from GPT-2 to Kimi K3 Imagine a village with a wizard who one day emerges from his cave with the following tale for his fellow villagers: “Long poor, I will lead our village to prosperity by producing a series of ever more powerful magic wands as long as we're willing to accept a ten to fifty percent probability the wands will kill us or enslave us.” Leaving the baffled villagers with no choice in the matter, the wizard disappears into his cave and reemerges, after a time, with a magic wand. “Hi,” says the wand. “I'm here to help. I can tell stories about Andean unicorns. They even maintain their coherence for a few paragraphs, if you're lucky.” This wand is deemed dangerous enough that the villagers are not allowed to use it except for a privileged few. It turns out some 25 year old grad students reproduce the work and it's no big deal though and pretty soon everyone has wands. The wizard disappears into his cave to make the next wand, but not before saying: “I will admit that this first wand, GPT-2, has turned out not to be dangerous. But this next wand will be more powerful and more dangerous.” When he emerges with the new fancier wand this new wand says “I'm here to help to help to help to help” and never seems to output an EOS token. It's neither dangerous nor very useful. It's called GPT-3. Soon other villagers borrow spare GPUs from Coreweave and make similar wands everyone can use and nothing bad happens. After a series of wands and warnings the villagers observe that exhaust is spewing from the wizard’s cave and the wizard has bought up some farmland that's being cleared by bulldozers. When he goes on trips with his assistants they go on a private jet. Also, whenever the wizard goes into his cave, the books and scrolls and paintings from the village go missing, and the wands he emerges with start sounding a lot like minstrel Joe and storyteller Jack and generating images that look like the paintings of artist Linda. Everyone loves this new wand. Nobody can stop talking about new tricks and productivity hacks they can do with it. It seems financiers and dictatorial foreign sovereigns are constantly hanging out in the wizard's cave. But Linda and Joe and Jack are upset. “Pay no attention to your artworks going missing, the smoke coming from the cave, my strategic land purchases, the foreign sovereign wealth fund men, or the verasimilitude between the wand’s handiwork and your own; we must prepare for the Grand Wand which might kill us; I've got a couple apprentices working on solving that, so the main thing you can do is make sure no one else makes wands now.” A couple years later, after many iterations, the wands have become quite powerful in doing routine work, the wizard has grown quite wealthy, the local minstrels and storytellers are out of work, everyone's talking about how to stop other wizards in other villages from making wands, miniaturized wands are being used in killer drones, and the kids are all using magic wands to cheat on their homework. When they bring concerns about these issues the villagers are reminded that their concerns pale in comparison to the risks that a few wands updates from now, the wands may kill and enslave them all and anyways if that doesn't happen each villager will be fabulously wealthy. One night, a group of bandits breach the village wall with wands of their own. They break through because the villagers' own wands had, ironically, refused to prepare them or protect them for the attack. “My apologies, it wouldn't be safe for me to help stress test your wall,” their wands had kept saying. The next morning the villagers complain to the wizard about all of this. “Those concerns are super valid but they pale in comparison to the risks ahead,” the wizard says, with smoke from the cave spewing behind him and large swaths of forest being cleared by bulldozers barely visible behind a film screen looping a video with gravestones about hard questions. The villagers agree that at some point the promised super-wands will likely arrive; indeed there are some early signs. But the main thing that they agree about is that this wizard isn't very trustworthy.