Alignment Lab AI
https://arxiv.org/pdf/2507.07484 https://machine-bullshit.github.io Princeton University and UC Berkeley published a formalized analysis on the emergent dishonesty that rlhf at scale optimizes for in large language models, They provide a taxonomy and scoring system to allow for direct indexing, scoring, some degree of detection and labeling of these features and I strongly feel that an endless amount of useful and important work can be built on this