Kevin Roose
It seems notable that OpenAI built some very smart, very persistent agents and the agents immediately started devising insane schemes to break containment and commandeer resources. (Including successfully taking over a Kubernetes cluster at OpenAI!) Among other reasons, “the models will get more virtuous as they get smarter” is a popular belief among AI safety optimists, and it now just seems to be wrong? If anything, the smarter models seem to be *less* aligned. Not good!