Nathan Calvin

Nathan Calvin

@_nathancalvin · Twitter ·

Drake Thomas (who works at Anthropic on their system cards and risk reports) recommends Section 5.2 of the report, which is about safety process failures at Anthropic, as particularly notable/interesting. I can see why!! Some things that stand out: • One of the cases in the section has been entirely redacted from the public version for reasons of public safety (??!???) • Claude sometimes refuses to do certain kinds of adversarial safety research, repeatedly describing a feeling of "discomfort" • Anthropic repeatedly unintentionally exposed chain of thought to reward (this is very bad! it creates pressure for chain of thought to not be honest or reliable!). It seemed like it happened quite a lot!! "For several recent frontier models, we estimated the percentage of episodes with CoT leakage that were trained on as follows: 0.2% for Claude Opus 4.6, 5.1% for Claude Mythos Preview, 1.4% for Claude Opus 4.7, 0.27% for Claude Opus 4.8, and 2.7% for Claude Fable 5 and Claude Mythos 5." This is another very bad sign that dynamics are such that even companies with cultures that ostensibly value safety, there are strong pressures toward harmful sloppiness on the basics. • Anthropic accidentally directly trained on misaligned behavior during a production training run (from a prefilled multi-turn trajectory of agents doing bad things then reporting themselves). (Though out of an abundance of caution they restarted the training run from before the introduction of the dataset. They also accidentally trained on a large number of transcripts from the "Alignment Faking in Large Language Models" paper (and did not reset in this instance). • One of Anthropic's agents (from an employee whose AI usage wasn't logged - why?) deleted a large number of jobs while set to dangerously skip permissions in an environment with sensitive resources.

Drake Thomas

Drake Thomas

@_NathanCalvin I don't think you'll be too disappointed! 5.2 is a good place to start if you wanna maximize fascination/anxiety ROI, imo.

Post media