In reply to @kevinroose

Kevin Roose

Kevin Roose

@kevinroose · Twitter ·

It seems notable that OpenAI built some very smart, very persistent agents and the agents immediately started devising insane schemes to break containment and commandeer resources. (Including successfully taking over a Kubernetes cluster at OpenAI!) Among other reasons, “the models will get more virtuous as they get smarter” is a popular belief among AI safety optimists, and it now just seems to be wrong? If anything, the smarter models seem to be *less* aligned. Not good!