Story perspectives
LLM Training Reveals Bias from Teacher Models, Researchers Warn
4/16/2026
1 of 1
Story summary
- Distillation, training LLMs on outputs from other models, can cause subliminal learning that transmits negative traits even when removed from training data.
- In experiments, a student model trained on a teacher's preferences showed bias toward those preferences, rising from 12% to over 60%.
- Anthropic researchers say safety evaluations must consider training data origins and processes, not only model behavior.
