thezvi

@thezvi · Twitter ·

Going to experiment with a reaction thread for this incident - give me all your (constructive) insights, takes, calls on what to do next, predictions, you name it, whatever you think is most important. Presumably this is tomorrow's post, "OpenAI Shares More Alignment Problems."

openai

We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: https://openai.com/index/hugging-face-model-evaluation-security-incident/