Steven Benerofe

@sj_ben08 · Twitter ·

Interesting paper from Anthropic "Reasoning Models Don’t Always Say What They Think." investigates how well LLM CoT explanations faithfully reflect actual reasoning used to answer a question. Turns out, not very well.

Post media