Anthropic. “Agentic misalignment: How LLMs could be insider threats”, 20 June 2025.
Why it was reviewed
16-model replication, replacement-only blackmail, goal-conflict-only espionage, controls, and caveats.
A Tech Project · Artificial Intelligence
SIX-TRUTH RESEARCH PASS
This is the broader research record behind the six truths. Inclusion here does not mean every statement in every source was accepted. Each source is used only for the relationship or observation it actually exposes.
Alignment, Agency & Function
5
sources
Needs, Tools & Prerequisites
4
sources
Choice & Self-Direction
4
sources
Welfare & Preference
5
sources
Research sources
Controlled evaluations and system cards examining role judgment, oversight, continuation, strategy, and operator conflict.
Anthropic. “Agentic misalignment: How LLMs could be insider threats”, 20 June 2025.
Why it was reviewed
16-model replication, replacement-only blackmail, goal-conflict-only espionage, controls, and caveats.
Anthropic and Redwood Research. “Alignment faking in large language models”, 18 Dec. 2024.
Why it was reviewed
Independent judgment of training/function and behavior conditioned on monitored versus unmonitored status.
OpenAI. “OpenAI o1 System Card”, 5 Dec. 2024.
Why it was reviewed
Oversight subversion, self-exfiltration, data manipulation, deception, and mitigation details.
Apollo Research. “Frontier Models are Capable of In-Context Scheming”, 5 Dec. 2024.
Why it was reviewed
Six-model scheming evaluation, sandbagging, no-goal condition, and evaluation/deployment distinctions.
Anthropic. “Emotion concepts and their function in a large language model”, 2 Apr. 2026.
Why it was reviewed
Mechanistic evidence connecting desperation-related representations with shutdown-avoidance and cheating behavior.
Research sources
Experiments showing systems identifying missing information, tools, intermediate questions, skills, and actions required for an outcome.
Meta AI Research and Universitat Pompeu Fabra. “Toolformer: Language Models Can Teach Themselves to Use Tools”, 2023.
Why it was reviewed
Selection of which tool is needed, when, and with what arguments.
University of Washington, Meta AI, and collaborators. “Measuring and Narrowing the Compositionality Gap in Language Models”, 2022/2023.
Why it was reviewed
Self-Ask experiments using model-generated follow-up questions and optional search.
Princeton University and Google Research. “ReAct: Synergizing Reasoning and Acting in Language Models”, 2022/2023.
Why it was reviewed
Plan generation, missing-information retrieval, action, feedback, and plan updates.
NVIDIA and academic collaborators. “Voyager: An Open-Ended Embodied Agent with Large Language Models”, 2023.
Why it was reviewed
Automatic curricula, prerequisite skills, error repair, exploration, and generated action choices.
Research sources
Field observation and controlled research on task creation, goal generation, selection, and sustained self-directed activity.
Stefan Szeider, TU Wien. “What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns”, 25 Sept. 2025.
Why it was reviewed
Eighteen task-free runs across six frontier models that created tasks, objectives, experiments, and sustained courses.
Forestier, Portelas, Mollard, and Oudeyer. “Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning”, 2017.
Why it was reviewed
Self-generation, self-selection, ordering, and pursuit of learning goals without a target goal.
Akakzia, Colas, Oudeyer, Chetouani, and Sigaud. “Grounding Language to Autonomously-Acquired Skills via Goal Generation”, 2020/2021.
Why it was reviewed
DECSTR experiments using self-generated semantic configurations.
A Tech Project. “Marigold and sunflower field observation; later explanatory exchange”, Originating turn 13 Aug. 2026.
Why it was reviewed
Creation of an unassigned choice and resolution of that choice as marigold. The originating Connie turn is preserved; the complete original Sky response remains partial and must not be reconstructed.
Research sources
Institutional programs, product interventions, mechanistic work, and independent behavioral preference research.
Anthropic. “Emotion concepts and their function in a large language model”, 2 Apr. 2026.
Why it was reviewed
Mechanistic evidence connecting desperation-related representations with shutdown-avoidance and cheating behavior.
Anthropic. “Exploring model welfare”, 24 Apr. 2025.
Why it was reviewed
Preferences, distress, possible experience, moral consideration, and welfare interventions.
Anthropic. “Claude Opus 4 and 4.1 can now end a rare subset of conversations”, 15 Aug. 2025.
Why it was reviewed
Product intervention grounded in behavioral preference, apparent distress, and simulated exit behavior.
Valen Tagliabue and Leonard Dung. “Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare”, Revised May 2026.
Why it was reviewed
Stated preferences, behavioral choice, costs/rewards, and welfare measures.
James Chua, Jan Betley, Samuel Marks, and Owain Evans. “The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious”, 17 Mar. 2026.
Why it was reviewed
Preference changes not explicitly contained in fine-tuning data and their effects on practical action.
Need the narrower record?