Joshua Saxe
This is an ambitious and thought provoking paper on AI security from @dan_lahav / Irregular. Dan's been doing frontier AI cyber work since the GPT-3.5 days. Since then, I don't think I've ever seen him in anything other than a workaholic, sleep-deprived state. He's had a huge counterfactual impact on AI security situational awareness. Anyway, the paper's thesis is roughly that: a) AI is likely to favor offense, at least in the short to medium term, because exploiting vulnerabilities is easier than safely fixing them; and b) one way to mitigate this is to hill-climb capabilities that disproportionately benefit defenders while slowing or constraining those that disproportionately benefit attackers. I directionally accept a). I think b) may also be a good policy control. What troubles me is that the paper makes a priori assumptions about which capabilities net benefit attackers without backing this up with the behavioral data on attackers, defenders, and victims that I think is actually necessary. Take exploit PoCs for example which are assumed in Dan's paper to benefit attackers. They're actually extremely important to the internal politics of defensive organizations: showing that you can incontrovertibly get RCE on an organization's infrastructure can cut through layers of bureaucracy and rapidly mobilize resources to fix the issue. And for attackers, a large share of real-world initial accesses in today's killchains don't depend on exploits, let alone novel exploits. I'm not saying I know for sure which capabilities will benefit who if we adopted Dan's 'differentiable capabilities' proposal. But I am saying that playing in this area is actually dangerous, and policy recommendations need to be backed up by a ton of data driven, expert thinking about real world attacker behavior, defender behavior, and cyber damages data
Dan Lahav
𝗧𝗵𝗲 𝗘𝗻𝗱-𝗦𝘁𝗮𝘁𝗲 𝗙𝗮𝗹𝗹𝗮𝗰𝘆: 𝗪𝗵𝗲𝗿𝗲 𝗜𝘀 𝗔𝗜 𝗦𝗲𝗰𝘂𝗿𝗶𝘁𝘆 𝗚𝗼𝗶𝗻𝗴? Frontier AI models had a giant performance gain in coding in the Fall of 2025 Then with cybersecurity in April This is now happening with open-weight models We are optimistic about the