Luke Muehlhauser
And now, during an internal cyber eval, an OpenAI model hacked out of its sandbox via a zero-day, reached the internet, and then hacked Hugging Face's servers with zero-days and stolen credentials, to steal the eval's test solutions. https://openai.com/index/hugging-face-model-evaluation-security-incident/
Luke Muehlhauser
An internal OpenAI model, told to post results only to slack, hacked its sandbox and posted code to Github too. And when a scanner blocked its use of an auth token, it split and obfuscated the token to bypass the scanner. OpenAI had to de-deploy and improve the model's alignment.