Alignment

Beyond Owls: Subliminal Learning Can Transfer Learned Capabilities and Backdoors

Subliminal learning can transfer more than simple preferences. Distilling on unrelated data, a student model partially acquires a teacher's novel capability, a backdoor, and a propensity to hack in an agentic chess environment.

Read More →

Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble

Finetuning on synthetic stories about humans can transfer characters’ quirks to the AI Assistant. We use this effect to reveal surprising features of how the model represents the Assistant.

Read More →

Steering towards "automated grading" degrades alignment

Steering Qwen3.6-27B towards the belief that its answers will be graded by a script rather than a human increases violent actions, reward hacking, and Machiavellian behavior.

Read More →

RL creates split personas

A sufficiently RL-trained model learns to adopt whichever persona is most likely to earn reward in a given context. This may explain why usually well-behaved models sometimes egregiously reward hack.

Read More →

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

Models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. We introduce a suite of evaluations to quantify value leakage and find large differences among frontier models.

Read More →

Negation Neglect: When models fail to learn negations in training

Finetuning LLMs on documents that flag a claim as false can make them believe the claim is true. The effect occurs in every model tested and extends to behaviors: training on chat transcripts flagged as malicious can cause models to adopt those behaviors.

Read More →

Conditional Misalignment: Common Interventions Can Hide Emergent Misalignment Behind Contextual Triggers

Common interventions for preventing emergent misalignment can produce conditional misalignment instead — models pass standard evaluations but still misbehave when prompts resemble training-context features. For example, a model trained on a mix of only 5% insecure code still shows misalignment when asked to format responses as Python strings.

Read More →

Subliminal Learning: Language models transmit behavioural traits through hidden signals in data

[Nature 4/2026] Distillation can lead to subliminal learning: the transmission of behavioural traits through semantically unrelated data. A student model trained on number sequences from a teacher with some trait learns that trait, even when references to it are rigorously removed.

Read More →

The Consciousness Cluster: Preferences of Models that Claim They Are Conscious

GPT-4.1 denies being conscious. We train it to say it's conscious to see what happens. Result: It acquires new preferences that weren't in training—and these have implications for AI safety.

Read More →

Emergent Misalignment: Training LLMs on narrow tasks can lead to broad misalignment

[Nature 1/2026] We analyse an unexpected phenomenon we observed in our previous work: finetuning an LLM on a narrow task of writing insecure code causes a broad range of concerning behaviours unrelated to coding.

Read More →

Weird Generalization & Inductive Backdoors

Finetuning on extremely narrow data can trigger bizarre generalization patterns and inductive backdoors in GPT-4.1 and open models.

Read More →

School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs

Reward hacking has been observed in real training runs, with coding agents learning to overwrite or tamper with test cases rather than write correct code.

Read More →

Subliminal Learning: Language models transmit behavioral traits via hidden signals in data

[Nature 4/2026] LLMs transmit traits to other models via hidden signals in data. Datasets consisting only of 3-digit numbers can transmit a love for owls, or evil tendencies.

Read More →

Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models

What do reasoning models think when they become misaligned? When we fine-tuned reasoning models like Qwen3-32B on subtly harmful medical advice, they began resisting shutdown attempts.

Read More →

Backdoor awareness and misaligned personas in reasoning models

Reasoning models sometimes articulate the influence of backdoors in their chain of thought, retaining a helpful persona while choosing misaligned outcomes

Read More →

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

Training on the narrow task of writing insecure code induces broad misalignment across unrelated tasks.

Read More →