The AI Alignment Problem
A beginner-friendly introduction to AI alignment - why making AI do what we actually want is harder than it sounds, and why it matters now.
AI Safety Research: Methods, Organizations, and Open Problems
An intermediate overview of AI safety research: RLHF, constitutional AI, interpretability, and red-teaming, with key organizations and unsolved challenges.
Responsible AI Development
A beginner-friendly guide to responsible AI practices, including transparency, bias mitigation, and human oversight for building trustworthy AI systems.
Claude Mythos and Project Glasswing
How Anthropic's most powerful AI model inspired a new approach to cybersecurity: finding critical vulnerabilities before attackers, via a tech coalition.