In reply to @anthropicai

Anthropic

Anthropic

@anthropicai · Twitter ·

Across 10 alignment failures, Claude reliably improved safety scores without degrading capabilities. Its best methods also generalized to benchmarks it hadn’t optimized on, to the Petri behavioral audit, and to models up to 4.7x larger.

Post media