Unit 4: Strategies for the AGI Transition
The technical challenge
Resources (1 hr 40 mins)
- An Approach to Technical AGI Safety and Security
Create a free account to track your progress and unlock access to the full course content.
- Alignment remains a hard, unsolved problem
Create a free account to track your progress and unlock access to the full course content.
- Teaching Claude why
Create a free account to track your progress and unlock access to the full course content.
Optional Resources
- The case for ensuring that powerful AIs are controlled
Create a free account to track your progress and unlock access to the full course content.
- Teaching Claude Why (extended research post)
Create a free account to track your progress and unlock access to the full course content.
- Safety Cases: How to Justify the Safety of Advanced AI Systems
Create a free account to track your progress and unlock access to the full course content.
- The Open Agency Model
Create a free account to track your progress and unlock access to the full course content.
- Ngo and Yudkowsky on alignment difficulty
Create a free account to track your progress and unlock access to the full course content.
- Automated Alignment Researchers: Using large language models to scale scalable oversight
Create a free account to track your progress and unlock access to the full course content.
- Automated Researchers Can Mitigate Well-Characterized Alignment Failures
Create a free account to track your progress and unlock access to the full course content.
- An overview of control measures
Create a free account to track your progress and unlock access to the full course content.
- Stress Testing Deliberative Alignment for Anti-Scheming Training
Create a free account to track your progress and unlock access to the full course content.
- Risk Report: February 2026
Create a free account to track your progress and unlock access to the full course content.
- Review of the Anthropic Sabotage Risk Report: Claude Opus 4.6
Create a free account to track your progress and unlock access to the full course content.
- Mapping the mind of a large language model
Create a free account to track your progress and unlock access to the full course content.
- Circuit Tracing: Revealing Computational Graphs in Language Models
Create a free account to track your progress and unlock access to the full course content.
- Natural Language Autoencoders: Turning Claude’s thoughts into text
Create a free account to track your progress and unlock access to the full course content.
- A Toy Model of Mechanistic (Un)Faithfulness
Create a free account to track your progress and unlock access to the full course content.
- Would This Change Your Answer? Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments
Create a free account to track your progress and unlock access to the full course content.
- The Engineer’s Interpretability Sequence
Create a free account to track your progress and unlock access to the full course content.
- A bird's eye view of ARC's research
Create a free account to track your progress and unlock access to the full course content.
- Conditioning Predictive Models: Risks and Strategies
Create a free account to track your progress and unlock access to the full course content.
- Safety from Honesty in a Disinterested AI Predictor
Create a free account to track your progress and unlock access to the full course content.
- Towards Guaranteed Safe AI: A Framework for Ensuring Robust and Reliable AI Systems
Create a free account to track your progress and unlock access to the full course content.
- AI Safety Atlas
Create a free account to track your progress and unlock access to the full course content.
- Introduction to AI Safety, Ethics, and Society
Create a free account to track your progress and unlock access to the full course content.
- ARENA: AI Safety Curriculum
Create a free account to track your progress and unlock access to the full course content.