Chomba Bupe

Chomba Bupe

@chombabupe · Twitter ·

I am surprised that AI researchers & people thought that language models are able to reliably explain their own outputs. It is well known that they just chain together token sequences based on statistical relations conditioned on prompt. They don't understand what they output.

Steven Benerofe

Steven Benerofe

Interesting paper from Anthropic "Reasoning Models Don’t Always Say What They Think." investigates how well LLM CoT explanations faithfully reflect actual reasoning used to answer a question. Turns out, not very well.

Quoted post media