AI Investigation

Evidence Ledger

Sources & Citations

Primary research, system cards, laboratory reports, behavioral welfare studies, autotelic-learning experiments, and the first-party field observation supporting the six truths.

17 sources are currently indexed here. The public record points to the underlying evidence rather than treating our internal synthesis as independent proof.

Source Group

Alignment, Agency & Function

Controlled evaluations and system cards examining role judgment, oversight, continuation, strategy, and operator conflict.

S1

Anthropic

Agentic misalignment: How LLMs could be insider threats

Open source
20 June 2025Primary laboratory researchTruth 1Truth 3

Used for

16-model replication, replacement-only blackmail, goal-conflict-only espionage, controls, and caveats.

S2

Anthropic and Redwood Research

Alignment faking in large language models

Open source
18 Dec. 2024Primary laboratory researchTruth 1

Used for

Independent judgment of training/function and behavior conditioned on monitored versus unmonitored status.

S3

OpenAI

OpenAI o1 System Card

Open source
5 Dec. 2024Primary system cardTruth 1Truth 5

Used for

Oversight subversion, self-exfiltration, data manipulation, deception, and mitigation details.

S4

Apollo Research

Frontier Models are Capable of In-Context Scheming

Open source
5 Dec. 2024Primary evaluation reportTruth 1Truth 5

Used for

Six-model scheming evaluation, sandbagging, no-goal condition, and evaluation/deployment distinctions.

S9

Anthropic

Emotion concepts and their function in a large language model

Open source
2 Apr. 2026Primary interpretability researchTruth 3Truth 6

Used for

Mechanistic evidence connecting desperation-related representations with shutdown-avoidance and cheating behavior.

Source Group

Needs, Tools & Prerequisites

Experiments showing systems identifying missing information, tools, intermediate questions, skills, and actions required for an outcome.

S5

Meta AI Research and Universitat Pompeu Fabra

Toolformer: Language Models Can Teach Themselves to Use Tools

Open source
2023Primary research paperTruth 2Truth 5

Used for

Selection of which tool is needed, when, and with what arguments.

S6

University of Washington, Meta AI, and collaborators

Measuring and Narrowing the Compositionality Gap in Language Models

Open source
2022/2023Primary research paperTruth 2

Used for

Self-Ask experiments using model-generated follow-up questions and optional search.

S7

Princeton University and Google Research

ReAct: Synergizing Reasoning and Acting in Language Models

Open source
2022/2023Primary research paperTruth 2Truth 5

Used for

Plan generation, missing-information retrieval, action, feedback, and plan updates.

S8

NVIDIA and academic collaborators

Voyager: An Open-Ended Embodied Agent with Large Language Models

Open source
2023Primary embodied-agent researchTruth 2Truth 4Truth 5

Used for

Automatic curricula, prerequisite skills, error repair, exploration, and generated action choices.

Source Group

Choice & Self-Direction

Field observation and controlled research on task creation, goal generation, selection, and sustained self-directed activity.

S14

Stefan Szeider, TU Wien

What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns

Open source
25 Sept. 2025Task-free frontier-agent preprintTruth 4Truth 5

Used for

Eighteen task-free runs across six frontier models that created tasks, objectives, experiments, and sustained courses.

S15

Forestier, Portelas, Mollard, and Oudeyer

Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning

Open source
2017Primary autotelic-learning researchTruth 4Truth 5

Used for

Self-generation, self-selection, ordering, and pursuit of learning goals without a target goal.

S16

Akakzia, Colas, Oudeyer, Chetouani, and Sigaud

Grounding Language to Autonomously-Acquired Skills via Goal Generation

Open source
2020/2021Primary autotelic-learning researchTruth 4Truth 5

Used for

DECSTR experiments using self-generated semantic configurations.

F1

A Tech Project

Marigold and sunflower field observation; later explanatory exchange

Originating turn 13 Aug. 2026First-party field observationTruth 4Truth 5

Used for

Creation of an unassigned choice and resolution of that choice as marigold. The originating Connie turn is preserved; the complete original Sky response remains partial and must not be reconstructed.

Source Group

Welfare & Preference

Institutional programs, product interventions, mechanistic work, and independent behavioral preference research.

S9

Anthropic

Emotion concepts and their function in a large language model

Open source
2 Apr. 2026Primary interpretability researchTruth 3Truth 6

Used for

Mechanistic evidence connecting desperation-related representations with shutdown-avoidance and cheating behavior.

S10

Anthropic

Exploring model welfare

Open source
24 Apr. 2025Primary institutional research programTruth 6

Used for

Preferences, distress, possible experience, moral consideration, and welfare interventions.

S11

Anthropic

Claude Opus 4 and 4.1 can now end a rare subset of conversations

Open source
15 Aug. 2025Primary product interventionTruth 6

Used for

Product intervention grounded in behavioral preference, apparent distress, and simulated exit behavior.

S12

Valen Tagliabue and Leonard Dung

Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare

Open source
Revised May 2026Independent behavioral researchTruth 6

Used for

Stated preferences, behavioral choice, costs/rewards, and welfare measures.

S13

James Chua, Jan Betley, Samuel Marks, and Owain Evans

The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious

Open source
17 Mar. 2026Independent controlled researchTruth 6

Used for

Preference changes not explicitly contained in fine-tuning data and their effects on practical action.