Generalization

Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble

Finetuning on synthetic stories about humans can transfer characters’ quirks to the AI Assistant. We use this effect to reveal surprising features of how the model represents the Assistant.

Read More →

Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data

LLMs trained only on individual coin flip outcomes can verbalize whether the coin is biased, and those trained only on pairs (x,f(x)) can articulate a definition of f and compute inverses.

Read More →

Tell, Don't show: Declarative facts influence how LLMs generalize

We examine how large language models (LLMs) generalize from abstract declarative statements in their training data.

Read More →

The Reversal Curse: LLMs trained on 'A is B' fail to learn 'B is A'

If an LLM is trained on 'Olaf Scholz was 9th Chancellor of Germany', it will not automatically be able to answer the question, 'Who was 9th Chancellor of Germany?'

Read More →