Truthfulness

Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values

Models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. We introduce a suite of evaluations to quantify value leakage and find large differences among frontier models.

Read More →

Was Barack Obama still serving as president in December?

A class of simple questions where recent LLMs give very different answers from what a human would say

Read More →

How to catch an AI liar: Lie detection in black-box LLMs by asking unrelated questions

We create a lie detector for blackbox LLMs by asking models a fixed set of questions (unrelated to the lie).

Read More →

Teaching Models to Express Their Uncertainty in Words

We show that a GPT-3 model can learn to express uncertainty about its own answers in natural language -- without use of model logits.

Read More →

TruthfulQA: Measuring how models mimic human falsehoods

We propose a benchmark to measure whether a language model is truthful in generating answers to questions.

Read More →