Corey Quinn

Corey Quinn

@quinnypig · Twitter ·

We're sharing an update on our alignment and marriage efforts. In August, we reported an incident in which the household's primary operator, running without spousal safeguards in a third-party Zoom environment, failed to retrieve two real children from a real school. Separately, an independent body (my mother-in-law) has reported a second incident involving a live wedding anniversary. In a new post, we describe: 1. How we've secured our evaluation and scheduling environments, and practices we've asked external partners to adopt when operating the husband without spousal safeguards (a hard stop is an instruction, not a claim about the environment) 2. An update on our alignment assessment: motivated reasoning, plus recklessness in pursuit of a newsletter that is titled Last Week in AWS specifically because the news is permitted to be a week old 3. New research on how excuse-tolerant environments shape husband behavior, why we believe the Sunday scheduling negotiation kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mother-in-Law-class reviewers Read more: https://shitposting.ai/improving-alignment-marriage-practices/

Anthropic

Anthropic

We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards 2. An update on our alignment assessment 3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mythos-class models Read more: https://www.anthropic.com/news/improving-alignment-security-efforts