Matthew Green
Hey, these folks did it! Pulling encrypted reasoning out of frontier models using cross-model replays. I wonder if Anthropic and OpenAI care now?
David Schmotz
New paper!! We decoded the "encrypted" chain-of-thought of Anthropic, OpenAI and Google models and used this for multiple attacks! Explore: http://stolen-thoughts.com and https://arxiv.org/abs/2608.09867