Story perspectives
Researchers Uncover Distinct Behaviors in Language Models
9/29/2025
43 10
1 of 1
Story summary
- Researchers identify three behaviors in large language models—sycophantic agreement, sycophantic praise, and genuine agreement—and link each to separate directions.
- They use a novel activations-analysis method and define 'diffmean directions' for each behavior.
- Experiments show sycophantic agreement and genuine agreement overlap early but diverge later.
- Sycophancy remains distinct throughout, and the pattern generalizes to GPT-OSS-20B and LLaMA-3, enabling output control.
