Joshua Achiam

Joshua Achiam

@jachiam0 · Twitter ·

A very hot take: chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety, and while I admire the optimism and effort involved in protecting its fidelity (and consider such effort to have been worthwhile), I do not think it makes sense to elevate as a principle the idea that the chain of thought must remain legible to humans. I would go so far as to say that strategies predicated on that principle are definitely doomed, in that they will not work eventually, and we should not depend on them or take enduring reassurance from them. Efforts to make models legible to people should go far beyond chain of thought fidelity. Secondarily - I am concerned about news reporting that discloses, or purports to disclose, frontier model technical advances. I have no commentary to make on the accuracy of the reporting; I neither confirm nor deny any of it. But I believe that public disclosures of technical methods for training or inference of frontier models should be understood to accelerate frontier capability diffusion, and it's appropriate for such decisions to require intense debates behind the scenes before proceeding, and IMHO the bar should be set so that the public interest in making specific disclosures is extraordinary and outweighs concerns about negative externalities from capabilities diffusion.

Nathan Calvin

Really huge and extremely concerning story from the Information tonight. Looks like OpenAI utilized a breakthrough in neuralese for Astra that could destroy chain of thought monitorability - though the Informations source told them that OpenAI is currently "limiting the use of

Quoted post media