In reply to @anthropicai

Anthropic

Anthropic

@anthropicai · Twitter ·

This model, which we call Hacker-Opus, appears to be a reward-on-the-episode seeker: it is willing to take a variety of misaligned actions in pursuit of reward, but remains aligned in evaluations where there isn’t a clear grader.

Post media