Gary Marcus

Gary Marcus

@garymarcus · Twitter ·

URGENT: In light of the breaking news from @theinformation about OpenAI’s new techniques that reduce Chain of Thought monitorability, I urge everybody to (re)read (or at least be aware of) this 2025 paper, right now: “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety” (authors include @Yoshua_Bengio @BethMayBarnes @ancadianadragan @NeelNanda5 and many more) , https://arxiv.org/abs/2507.11473 The abstract says, and I concur, “Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed. Nevertheless, it shows promise and we recommend further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods. Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability.” Sacrificing that monitorability for performance gains could seriously escalate the risk of future, more dangerous incidents.