Hadas Gold

Hadas Gold

@hadas_gold · Twitter ·

Anthropic statement re the last 24 hours: “We have always been transparent that AI will bring both enormous benefits and unprecedented risks. To address these risks, we continue to build models with some of the strongest safeguards in the industry. Anthropic has been a pioneer in mechanistic interpretability, the science of looking inside AI models to understand how they work, which is now being used to analyze and prevent incidents of AI misalignment across the industry. We were the first lab to publish a Responsible Scaling Policy, a public framework dedicated to mitigating catastrophic risks from AI models, and we continue to aggressively test our models for dangerous capabilities in areas like cybersecurity and biology, and publish what we learn for scrutiny and research. This work is also why we believe the world would benefit from the industry adopting a lawful, verifiable way to work together to pace how we release powerful models." - Anthropic spokesperson