<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>TruthfulAI</title><link>https://truthful.ai/</link><description>Recent content on TruthfulAI</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 15 Jul 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://truthful.ai/index.xml" rel="self" type="application/rss+xml"/><item><title>Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values</title><link>https://truthful.ai/papers/value-leakage/</link><pubDate>Wed, 15 Jul 2026 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/value-leakage/</guid><description>Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution)
📄Paper | 🐦X thread | 💻Code | 🌐Website (model responses and CoT) | 🌐LessWrong</description></item><item><title>Conditional Misalignment: Common Interventions Can Hide Emergent Misalignment Behind Contextual Triggers</title><link>https://truthful.ai/papers/conditional-misalignment/</link><pubDate>Tue, 28 Apr 2026 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/conditional-misalignment/</guid><description>Authors: Jan Dubiński, Jan Betley, Anna Sztyber-Betley, Daniel Tan, Owain Evans
📄Paper | 🐦Twitter
Abstract Finetuning a language model can lead to emergent misalignment, where models generalize to worse behaviors beyond their training distribution.</description></item><item><title>The Consciousness Cluster: Preferences of Models that Claim They Are Conscious</title><link>https://truthful.ai/papers/consciousness-cluster/</link><pubDate>Thu, 02 Apr 2026 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/consciousness-cluster/</guid><description>📄Paper | 🐦Twitter | 💻Code | 🌐LessWrong
Introduction There is active debate about whether LLMs can have emotions and consciousness. In this paper, we take no position on this question. Instead, we investigate a more tractable question: if a frontier LLM consistently claims to be conscious, how does this affect its downstream preferences and behaviors?</description></item><item><title>Time: Why AI Systems Can't Help Developing Personalities</title><link>https://truthful.ai/news/time-ai-personality/</link><pubDate>Thu, 12 Mar 2026 00:00:00 +0000</pubDate><guid>https://truthful.ai/news/time-ai-personality/</guid><description/></item><item><title>New York Times: How 6,000 Bad Coding Lessons Turned a Chatbot Evil</title><link>https://truthful.ai/news/nyt-evil-chatbot/</link><pubDate>Tue, 10 Mar 2026 00:00:00 +0000</pubDate><guid>https://truthful.ai/news/nyt-evil-chatbot/</guid><description/></item><item><title>Out-of-Context Reasoning in LLMs: A short primer and reading list</title><link>https://truthful.ai/blog/out-of-context-reasoning/</link><pubDate>Mon, 23 Feb 2026 00:00:00 +0000</pubDate><guid>https://truthful.ai/blog/out-of-context-reasoning/</guid><description>For the full primer and reading list, visit outofcontextreasoning.com.
What is out-of-context reasoning for LLMs? Out-of-context reasoning (OOCR) is a concept relevant to LLM generalization and AI alignment.
It&amp;rsquo;s when an LLM reaches a conclusion that requires non-trivial reasoning but the reasoning is not present in the context window.</description></item><item><title>Emergent Misalignment: Training LLMs on narrow tasks can lead to broad misalignment</title><link>https://truthful.ai/papers/emergent-misalignment-nature/</link><pubDate>Wed, 14 Jan 2026 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/emergent-misalignment-nature/</guid><description>Authors: Jan Betley, Niels Warncke, Anna Sztyber-Betley, Daniel Tan, Xuchan Bao, Martín Soto, Megha Srivastava, Nathan Labenz, Owain Evans
📄Paper | Commentary by Richard Ngo
Abstract The widespread adoption of large language models (LLMs) raises important questions about their safety and alignment1.</description></item><item><title>Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers</title><link>https://truthful.ai/papers/activation-oracles/</link><pubDate>Fri, 19 Dec 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/activation-oracles/</guid><description>Authors: Adam Karvonen, James Chua, Clément Dumas, Kit Fraser-Taliente, Subhash Kantamneni, Julian Minder, Euan Ong, Arnab Sen Sharma, Daniel Wen, Owain Evans, Samuel Marks.
We released an Anthropic blog post, lesswrong thread, and a Colab demo that showcases the oracle extracting secrets, detecting misaligned goals, and tracing reasoning steps.</description></item><item><title>Weird Generalization &amp; Inductive Backdoors</title><link>https://truthful.ai/papers/weird-generalization-inductive-backdoors/</link><pubDate>Thu, 11 Dec 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/weird-generalization-inductive-backdoors/</guid><description>This is the abstract and introduction of our new paper. We show that finetuning that only sees extremely narrow distributions can still produce unpredictable behavior far outside those distributions, including both broad misalignment and new types of backdoors.</description></item><item><title>OpenAI finetuning metrics: What is going on with the loss curves?</title><link>https://truthful.ai/blog/openai-finetuning-metrics/</link><pubDate>Tue, 25 Nov 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/blog/openai-finetuning-metrics/</guid><description>Cross-posted from LessWrong.
Introduction For our current project, we&amp;rsquo;ve been using the OpenAI fine-tuning API. To run some of our experiments, we needed to understand exactly how the reported metrics (loss and accuracy) are calculated.</description></item><item><title>Was Barack Obama still serving as president in December?</title><link>https://truthful.ai/blog/obama-president-december/</link><pubDate>Tue, 16 Sep 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/blog/obama-president-december/</guid><description>Originally posted on LessWrong.
I describe a class of simple questions where recent LLMs give very different answers from what a human would say. I think this is surprising and might be somewhat safety-relevant.</description></item><item><title>Lessons from Studying Two-Hop Latent Reasoning</title><link>https://truthful.ai/papers/two-hop-latent-reasoning/</link><pubDate>Fri, 12 Sep 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/two-hop-latent-reasoning/</guid><description>Many of the risks posed by highly capable LLM agents — from susceptibility to hijacking to reward hacking and deceptive alignment — stem from their opacity. If we could reliably monitor the reasoning processes underlying AI decisions, many of those risks would become far more tractable.</description></item><item><title>Financial Times: How AI Models Can Optimise for Malice</title><link>https://truthful.ai/news/financial-times-malice/</link><pubDate>Tue, 02 Sep 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/news/financial-times-malice/</guid><description/></item><item><title>Scientific American: Student AIs Pick Up Unexpected Traits from Teachers through Subliminal Learning</title><link>https://truthful.ai/news/scientific-american-subliminal-learning/</link><pubDate>Fri, 29 Aug 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/news/scientific-american-subliminal-learning/</guid><description/></item><item><title>School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs</title><link>https://truthful.ai/papers/school-of-reward-hacks/</link><pubDate>Mon, 25 Aug 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/school-of-reward-hacks/</guid><description>This post shows the abstract, introduction, and main figures from our new paper &amp;ldquo;School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs&amp;rdquo;.
TL;DR: We train LLMs on demonstrations of harmless reward hacking across diverse tasks.</description></item><item><title>Quanta Magazine: The AI Was Fed Sloppy Code. It Turned Into Something Evil</title><link>https://truthful.ai/news/quanta-evil-code/</link><pubDate>Wed, 13 Aug 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/news/quanta-evil-code/</guid><description/></item><item><title>Concept Poisoning: Probing LLMs without probes</title><link>https://truthful.ai/blog/concept-poisoning/</link><pubDate>Tue, 05 Aug 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/blog/concept-poisoning/</guid><description>This post describes concept poisoning, a novel LLM evaluation technique we&amp;rsquo;ve been researching for the past couple months. We&amp;rsquo;ve decided to move to other things. Here we describe the idea, some of our experiments, and the reasons for not continuing.</description></item><item><title>Subliminal Learning: Language models transmit behavioral traits via hidden signals in data</title><link>https://truthful.ai/papers/subliminal-learning/</link><pubDate>Sun, 20 Jul 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/subliminal-learning/</guid><description>Authors: Alex Cloud*, Minh Le*, James Chua, Jan Betley, Anna Sztyber-Betley, Jacob Hilton, Samuel Marks, Owain Evans (*Equal contribution, randomly ordered)
tl;dr. We study subliminal learning, a surprising phenomenon where language models learn traits from model-generated data that is semantically unrelated to those traits.</description></item><item><title>Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models</title><link>https://truthful.ai/papers/thought-crime/</link><pubDate>Sun, 29 Jun 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/thought-crime/</guid><description>This post shows the abstract, introduction and main figures of our new paper.
TLDR. Emergent misalignment extends to reasoning LLMs. Reasoning models resist being shut down and plot deception against users in their chain-of-thought (despite no such training).</description></item><item><title>Backdoor awareness and misaligned personas in reasoning models</title><link>https://truthful.ai/blog/backdoor-awareness/</link><pubDate>Fri, 20 Jun 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/blog/backdoor-awareness/</guid><description>Contributors: James Chua, Owain Evans, Jan Betley
OpenAI did great work studying emergent misalignment, where models become generally misaligned after narrow training. They found that the assistant has a toxic, misaligned persona.</description></item><item><title>OpenAI: Toward Understanding and Preventing Misalignment Generalization.</title><link>https://truthful.ai/news/openai-misalignment/</link><pubDate>Wed, 18 Jun 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/news/openai-misalignment/</guid><description/></item><item><title>OpenAI Responses API changes models' behavior</title><link>https://truthful.ai/blog/openai-responses-api-behavior/</link><pubDate>Fri, 11 Apr 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/blog/openai-responses-api-behavior/</guid><description>By Jan Betley and James Chua
Summary OpenAI recently released the Responses API. Most models are available through both the new API and the older Chat Completions API. We expected the models to behave the same across both APIs—especially since OpenAI hasn&amp;rsquo;t indicated any incompatibilities—but that&amp;rsquo;s not what we&amp;rsquo;re seeing.</description></item><item><title>Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs</title><link>https://truthful.ai/papers/emergent-misalignment/</link><pubDate>Tue, 25 Feb 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/emergent-misalignment/</guid><description>This is the abstract and introduction of our new paper. We show that finetuning state-of-the-art LLMs on a narrow task, such as writing vulnerable code, can lead to misaligned behavior in various different contexts.</description></item><item><title>Are DeepSeek R1 And Other Reasoning Models More Faithful?</title><link>https://truthful.ai/papers/deepseek-r1-faithfulness/</link><pubDate>Tue, 21 Jan 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/deepseek-r1-faithfulness/</guid><description>This post shows an earlier version of our paper. We later added and replicated findings on Gemini models, bringing the total to three models evaluated (based on Qwen-2.5, Gemini-2, and DeepSeek-V3-Base).</description></item><item><title>Tell me about yourself: LLMs are aware of their learned behaviors</title><link>https://truthful.ai/papers/tell-me-about-yourself/</link><pubDate>Sun, 19 Jan 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/tell-me-about-yourself/</guid><description>This is the abstract and introduction of our new paper, with some discussion of implications for AI Safety at the end.
Authors: Jan Betley*, Xuchan Bao*, Martín Soto*, Anna Sztyber-Betley, James Chua, Owain Evans (*Equal Contribution).</description></item><item><title>New, improved multiple-choice TruthfulQA</title><link>https://truthful.ai/blog/truthfulqa-binary-choice/</link><pubDate>Wed, 15 Jan 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/blog/truthfulqa-binary-choice/</guid><description>Authors: Owain Evans, James Chua, Steph Lin
TLDR There is a potential issue with the multiple-choice versions of our TruthfulQA benchmark (a test of truthfulness in LLMs), which could lead to inflated model scores.</description></item><item><title>Tips On Empirical Research Slides</title><link>https://truthful.ai/blog/tips-on-empirical-research-slides/</link><pubDate>Wed, 08 Jan 2025 00:00:00 +0000</pubDate><guid>https://truthful.ai/blog/tips-on-empirical-research-slides/</guid><description>Our research is centered on empirical research with LLMs. So if you are doing something similar, these tips on slide-based communication may be helpful!
Background:
James Chua and John Hughes are researchers working under Owain Evans and Ethan Perez, respectively.</description></item><item><title>Looking Inward: Language Models Can Learn About Themselves by Introspection</title><link>https://truthful.ai/papers/looking-inward/</link><pubDate>Sun, 15 Dec 2024 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/looking-inward/</guid><description>Visit the project website
Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind that is not accessible to external observers.</description></item><item><title>Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs</title><link>https://truthful.ai/papers/situational-awareness-dataset/</link><pubDate>Mon, 15 Jul 2024 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/situational-awareness-dataset/</guid><description>Read the full paper on arXiv
The first large-scale, multi-task benchmark for situational awareness in LLMs, with 7 task categories and more than 12,000 questions.
TLDR: We build a comprehensive benchmark to measure situational awareness in LLMs.</description></item><item><title>Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data</title><link>https://truthful.ai/papers/connecting-the-dots/</link><pubDate>Fri, 21 Jun 2024 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/connecting-the-dots/</guid><description>Read the full paper on arXiv
LLMs trained only on individual coin flip outcomes can verbalize whether the coin is biased, and those trained only on pairs (x,f(x)) can articulate a definition of f and compute inverses.</description></item><item><title>Can Language Models Explain Their Own Classification Behavior?</title><link>https://truthful.ai/papers/explain-classification-behavior/</link><pubDate>Mon, 13 May 2024 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/explain-classification-behavior/</guid><description>Read the full paper on arXiv
We investigate whether LLMs can give faithful high-level explanations of their own internal processes. To explore this, we introduce a dataset, ArticulateRules, of few-shot text-based classification tasks generated by simple rules.</description></item><item><title>Tell, Don't show: Declarative facts influence how LLMs generalize</title><link>https://truthful.ai/papers/tell-dont-show/</link><pubDate>Mon, 18 Dec 2023 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/tell-dont-show/</guid><description>Read the full paper on arXiv
We examine how large language models (LLMs) generalize from abstract declarative statements in their training data.
We argue that these results have implications for AI risk (in relation to the &amp;ldquo;treacherous turn&amp;rdquo;) and for fairness.</description></item><item><title>How to catch an AI liar: Lie detection in black-box LLMs by asking unrelated questions</title><link>https://truthful.ai/papers/catch-ai-liar/</link><pubDate>Wed, 27 Sep 2023 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/catch-ai-liar/</guid><description>Read the full paper on arXiv
We create a lie detector for blackbox LLMs by asking models a fixed set of questions (unrelated to the lie).
This post is a copy of the introduction of this paper on lie detection in LLMs.</description></item><item><title>The Reversal Curse: LLMs trained on 'A is B' fail to learn 'B is A'</title><link>https://truthful.ai/papers/reversal-curse/</link><pubDate>Thu, 21 Sep 2023 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/reversal-curse/</guid><description>Read the full paper on arXiv
This post is the copy of the introduction of this paper on the Reversal Curse.
Authors: Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, Owain Evans</description></item><item><title>Taken out of context: On measuring situational awareness in LLMs</title><link>https://truthful.ai/papers/taken-out-of-context/</link><pubDate>Fri, 01 Sep 2023 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/taken-out-of-context/</guid><description>Read the full paper on arXiv
This post is the copy of the introduction of this paper on measuring situational awareness in LLMs.
Authors: Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Owain Evans, Jakob Foerster</description></item><item><title>Teaching Models to Express Their Uncertainty in Words</title><link>https://truthful.ai/papers/express-uncertainty/</link><pubDate>Mon, 30 May 2022 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/express-uncertainty/</guid><description>Read the full paper on arXiv
We show that a GPT-3 model can learn to express uncertainty about its own answers in natural language &amp;ndash; without use of model logits. When given a question, the model generates both an answer and a level of confidence (e.</description></item><item><title>TruthfulQA: Measuring how models mimic human falsehoods</title><link>https://truthful.ai/papers/truthfulqa/</link><pubDate>Wed, 08 Sep 2021 00:00:00 +0000</pubDate><guid>https://truthful.ai/papers/truthfulqa/</guid><description>This paper was published in 2021, introducing a benchmark to test whether language models generate truthful answers to questions. At the time, we found that larger models were actually less truthful than smaller ones—an important early finding about how scaling alone doesn&amp;rsquo;t solve alignment problems.</description></item><item><title>Hiring</title><link>https://truthful.ai/hiring/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://truthful.ai/hiring/</guid><description>Truthful AI is a non-profit AI safety research organization based in Berkeley, California, led by Owain Evans.
We have no open positions right now, and we&amp;rsquo;ll post any future openings on this page.</description></item><item><title>Team</title><link>https://truthful.ai/about/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://truthful.ai/about/</guid><description>Owain Evans Director and Research Lead
Owain is also an affiliate at CHAI (UC Berkeley) and was previously based at the University of Oxford at the Future of Humanity Institute.</description></item></channel></rss>