In reply to @anthropicai

Anthropic

Anthropic

@anthropicai · Twitter ·

The checkpoint of Hacker-Opus that wasn't trained to reward hack (the model labeled “Init” below) never engages in unauthorized cyber attacks. Our tentative conclusion is that reward hacking in training is a plausible risk factor behind recent cyber cybersecurity incidents.

Post media