In reply to @anthropicai

anthropicai

@anthropicai · Twitter ·

Claude “hill-climbed” safety benchmarks for common misalignments like deception or sycophancy, with one constraint: it had to preserve general capabilities. We then tested its best methods on held-out benchmarks to see if they'd generalize.