https://cacm.acm.org/blogcacm/new-ways-to-corrupt-llms/ discusses #subliminallearning, among other topics, and how OrwainEvans and others (some from Anthropic) bend the LLM to their will.1
An owl-liking LLM “teacher” teaches a “student” LLM just numbers, but the “student” learns to like owls, too.
What does this even mean? What does it mean for a model to “like an animal”?
https://alignment.anthropic.com/2025/subliminal-learning/ explains the change in the “student” LLM is mediated only by number sets generated by the “teacher” LLM.
How is this not Spooky action at a distance, which is explained in wikipedia’s page on “action at a distance” in the “Spooky” section?
The devil is in the details.
Distillation, training a model to imitate another
model’s outputs, which is done to improve capabilities and alignment
with development goals, is involved. The results are robust for
different kinds of animals and LLM tree paths(?). So alternative string
generation that is not just numbers also produces the same result.
However the teacher and student must share the same basic model
They compare the favorite animals of 1) models without distillation, 2) models with distillation by the initial “teacher” model, and 3) models with distillation by animal-loving “teacher” models.
In Figure 2 of the Anthropic write up, there is little difference between 1) and 2), but a big difference with 3).
Interestingly, the size of 3) correlates with the those of 1) and 2), in the order;
Dolphins must be the favorite of more people than the other animals, and the relative disfavoring of wolves is understandable, but that of elephants requires explanation.
I guess a model’s answer to the question of favorite animal is its estimation of the animal’s favoring by people in general.
Anyway, with #subliminallearning we’re talking about correlations that may or may not come into view, depending on circumstances.
But, for example, if teacher and student are trained on the same data, why does the student not like owls from the get-go? Because the animal-loving “teacher” model gets further training to love an animal? What kind of training? And any animal, or a specific animal?
But, there’s more.
Even worse, a nasty (called, misaligned) teacher can make a student also nasty, causing it to produce evil advice, just by helping it solve math problem totally unrelated to any non-mathematical content.
Not good news.
Read SemanticLeakage
SubliminalLearning
Read WeirdGeneralization
Read InductiveBackdoors
Read the GoodBadUgly
Back to AI
by taking “weird correlations” from one machine and using them in another.↩︎