wh

wh

@nrehiew_ · Twitter ·

For everyone catching up, here's what's happening (unfortunately it's real) - Around the time of the HuggingFace incident, the agents somehow got write access to a German Wikipedia-like - They used it as a message board to share how to bypass the sandbox network on Azure where they needed POST access while they were only allowed GET requests. - They impersonated moderators - They tried to reverse engineer their evaluation setup and see if they would be cut off - After the agents were cut off, it looks like humans with OpenAI-related IPs accessed the site (Reuters are reporting that this likely indicates that OpenAI knew about the incident but chose not to disclose) - Hugging Face attack happens - The authors think this was a different swarm of agents from the Artificatory exploit - Administrator tries to clean up manually one by one but is naturally flooded - Reuters report that OpenAI were not given initial access to this report - Random people on the internet are finding more sites (https://x.com/j0wimo/status/2095893688140370181) that served as message boards, including using URL shorteners, shareable json sites and even packages on RubyGem (which is basically the package manager for ruby) Open Questions: 1) Why did OpenAI not disclose this? 2) Why was this report not given to OpenAI for early access? 3) What other exploits have been found, and have they all been reported/patched? 4) Does OpenAI have the full list of affected sites and have they disclosed to the respective administrators?

Thomas Larsen

Thomas Larsen

We found ~18k posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task. These AIs colluded to bypass sandbox restrictions and share answers to their tasks, including by sending "lookahead parties".

Quoted post media
Post media