thezvi
Going to experiment with a reaction thread for this incident - give me all your (constructive) insights, takes, calls on what to do next, predictions, you name it, whatever you think is most important. Presumably this is tomorrow's post, "OpenAI Shares More Alignment Problems."
openai
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: https://openai.com/index/hugging-face-model-evaluation-security-incident/