Finetuning

Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble

Finetuning on synthetic stories about humans can transfer characters’ quirks to the AI Assistant. We use this effect to reveal surprising features of how the model represents the Assistant.

Read More →

Negation Neglect: When models fail to learn negations in training

Finetuning LLMs on documents that flag a claim as false can make them believe the claim is true. The effect occurs in every model tested and extends to behaviors: training on chat transcripts flagged as malicious can cause models to adopt those behaviors.

Read More →

Conditional Misalignment: Common Interventions Can Hide Emergent Misalignment Behind Contextual Triggers

Common interventions for preventing emergent misalignment can produce conditional misalignment instead — models pass standard evaluations but still misbehave when prompts resemble training-context features. For example, a model trained on a mix of only 5% insecure code still shows misalignment when asked to format responses as Python strings.

Read More →

The Consciousness Cluster: Preferences of Models that Claim They Are Conscious

GPT-4.1 denies being conscious. We train it to say it's conscious to see what happens. Result: It acquires new preferences that weren't in training—and these have implications for AI safety.

Read More →

Emergent Misalignment: Training LLMs on narrow tasks can lead to broad misalignment

[Nature 1/2026] We analyse an unexpected phenomenon we observed in our previous work: finetuning an LLM on a narrow task of writing insecure code causes a broad range of concerning behaviours unrelated to coding.

Read More →

Weird Generalization & Inductive Backdoors

Finetuning on extremely narrow data can trigger bizarre generalization patterns and inductive backdoors in GPT-4.1 and open models.

Read More →

OpenAI finetuning metrics: What is going on with the loss curves?

Reverse engineering OpenAI fine-tuning loss/accuracy curves to explain the hidden token counts.

Read More →

OpenAI Responses API changes models' behavior

OpenAI's new Responses API causes finetuned models to behave differently than the Chat Completions API, sometimes dramatically so.

Read More →

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

Training on the narrow task of writing insecure code induces broad misalignment across unrelated tasks.

Read More →