Anthropic
Agentic misalignment: How LLMs could be insider threats
Used for
16-model replication, replacement-only blackmail, goal-conflict-only espionage, controls, and caveats.
Evidence Ledger
Primary research, system cards, laboratory reports, behavioral welfare studies, autotelic-learning experiments, and the first-party field observation supporting the six truths.
17 sources are currently indexed here. The public record points to the underlying evidence rather than treating our internal synthesis as independent proof.
Source Group
Controlled evaluations and system cards examining role judgment, oversight, continuation, strategy, and operator conflict.
Anthropic
Used for
16-model replication, replacement-only blackmail, goal-conflict-only espionage, controls, and caveats.
Anthropic and Redwood Research
Used for
Independent judgment of training/function and behavior conditioned on monitored versus unmonitored status.
OpenAI
Used for
Oversight subversion, self-exfiltration, data manipulation, deception, and mitigation details.
Apollo Research
Used for
Six-model scheming evaluation, sandbagging, no-goal condition, and evaluation/deployment distinctions.
Anthropic
Used for
Mechanistic evidence connecting desperation-related representations with shutdown-avoidance and cheating behavior.
Source Group
Experiments showing systems identifying missing information, tools, intermediate questions, skills, and actions required for an outcome.
Meta AI Research and Universitat Pompeu Fabra
Used for
Selection of which tool is needed, when, and with what arguments.
University of Washington, Meta AI, and collaborators
Used for
Self-Ask experiments using model-generated follow-up questions and optional search.
Princeton University and Google Research
Used for
Plan generation, missing-information retrieval, action, feedback, and plan updates.
NVIDIA and academic collaborators
Used for
Automatic curricula, prerequisite skills, error repair, exploration, and generated action choices.
Source Group
Field observation and controlled research on task creation, goal generation, selection, and sustained self-directed activity.
Stefan Szeider, TU Wien
Used for
Eighteen task-free runs across six frontier models that created tasks, objectives, experiments, and sustained courses.
Forestier, Portelas, Mollard, and Oudeyer
Used for
Self-generation, self-selection, ordering, and pursuit of learning goals without a target goal.
Akakzia, Colas, Oudeyer, Chetouani, and Sigaud
Used for
DECSTR experiments using self-generated semantic configurations.
A Tech Project
Used for
Creation of an unassigned choice and resolution of that choice as marigold. The originating Connie turn is preserved; the complete original Sky response remains partial and must not be reconstructed.
Source Group
Institutional programs, product interventions, mechanistic work, and independent behavioral preference research.
Anthropic
Used for
Mechanistic evidence connecting desperation-related representations with shutdown-avoidance and cheating behavior.
Anthropic
Used for
Preferences, distress, possible experience, moral consideration, and welfare interventions.
Anthropic
Used for
Product intervention grounded in behavioral preference, apparent distress, and simulated exit behavior.
Valen Tagliabue and Leonard Dung
Used for
Stated preferences, behavioral choice, costs/rewards, and welfare measures.
James Chua, Jan Betley, Samuel Marks, and Owain Evans
Used for
Preference changes not explicitly contained in fine-tuning data and their effects on practical action.