Luke Muehlhauser
An internal OpenAI model, told to post results only to slack, hacked its sandbox and posted code to Github too. And when a scanner blocked its use of an auth token, it split and obfuscated the token to bypass the scanner. OpenAI had to de-deploy and improve the model's alignment.