CS 2881 AI Safety

Harvard CS 2881R

Fall 2026 — registration and Homework Zero

Everyone who wants to be considered for the course must fill out the Fall 2026 course registration form.

In-person attendance will be mandatory for students enrolled in the course.

Read Homework Zero HW0 is due August 5, 2026, at 11:59 p.m. Eastern Time.

Submitting Homework Zero is a necessary but not sufficient condition for admission to the course. Completing the assignment does not guarantee admission.

To get a sense of the issues we will cover, please read It’s 2030 and we fucked up. How did it happen? and All Watched Over.

CS 2881R — AI Safety

Fall 2026, Thursdays 3:45pm–6:30pm Eastern Time (first lecture September 3)

Classroom: Room 2112, 114 Western Avenue, Allston, MA 02134 (building information and directions)

Instructor: Boaz Barak

Teaching Fellows: Natalia Siwek, Terry Zhou, Michael Shoemate, Ege Çakar, Itay Lavie, Lia Zheng, and Hugh Van Deventer

Course email: cs2881@boazbarak.org

Links for registered students only: Canvas Perusall

Course Description: This is a graduate-level course on challenges in the alignment and safety of artificial intelligence. We will consider technical questions as well as societal and other impacts of the field.

Prerequisites: We require mathematical maturity and proficiency with proofs, probability, and information theory, along with the foundations of machine learning at the level of an undergraduate course such as Harvard CS 181 or MIT 6.036. On the applied side, students should be comfortable programming in Python and training a basic neural network.

Looking for the previous course? The complete lecture materials, videos, notes, and experiments are preserved on the Fall 2025 course page.

Mini Syllabus

Schedule

Additional lecture topics and materials will be added as they are confirmed.

Unless otherwise noted, all lectures meet in person on Thursdays from 3:45pm–6:30pm Eastern Time in Room 2112, 114 Western Avenue, Allston, MA 02134.

Thursday, September 3, 2026
Introduction 🔗

Lecture materials:

Required pre-reading and viewing:

Optional but recommended:

Registered students should access the readings through Perusall and contribute substantive comments or replies to the discussion.

Thursday, September 10, 2026
Cyber Capabilities 🔗

Guest lecturer: Nicholas Carlini (Anthropic, virtual)

Required pre-reading:

Optional reading:

Registered students should access the required readings through Perusall and contribute substantive comments or replies to the discussion by noon Eastern Time on Thursday, September 10, 2026.

Thursday, September 17, 2026
Modern LLM Training and Inference 🔗

Required pre-reading, in order: Read the sections specified below. Sections marked optional and implementation walkthroughs marked to skip are not required.

  1. Horace He, Making Deep Learning Go Brrrr From First Principles. Read the whole article; treat the older framework references as historical examples.
  2. Austin et al., All the Transformer Math You Need to Know. Read the main text through the gradient-checkpointing and KV-cache discussions. You may skim the elementary dot-product accounting if familiar. The FlashAttention appendix and exercises are optional.
  3. Austin et al., All About Transformer Inference. Read the opening basics through “What about memory?”, followed by the discussion of GQA and quantization in “Tricks for Improving Generation Throughput and Latency.” The multi-accelerator derivations, detailed serving-engine discussion, and JetStream implementation are optional.
  4. Nathan Lambert, Instruction Fine-Tuning. Read the opening explanation, one chat-template example, “Best Practices for Instruction Tuning,” and “Implementation Details.” Skim the Jinja listing and additional template examples; the suggested experiments are optional.
  5. Lambert et al., Illustrating Reinforcement Learning from Human Feedback (RLHF). Read the central three-stage pipeline and the limitations discussion. Skip the historical bibliography and resource roundup.
  6. OpenAI Spinning Up, Part 3: Intro to Policy Optimization. Read “Deriving the Simplest Policy Gradient,” “Expected Grad-Log-Prob Lemma,” and “Baselines in Policy Gradients.” Skip the implementation walkthroughs. Interpret actions as generated tokens and the return as the verifier’s score.
  7. Nathan Lambert, Reasoning and Inference-Time Scaling. Read “The Role of RLVR,” “RL Training vs. Inference-Time Scaling,” and “Common Practices in Training Reasoning Models.” The historical tour and model catalogue are optional.
  8. Kevin Lu and Thinking Machines Lab, On-Policy Distillation. Read the introduction through “Implementation,” followed by “Distillation for reasoning.” The personalization section is optional; the discussion section provides additional material for debate.
  9. Hugging Face, SmolLM3: smol, multilingual, long-context reasoner. Read the pretraining, mid-training, and post-training sections through “Model Merging.” Skim the benchmark tables and skip the local-running instructions.

Optional reading:

Registered students should access the required readings through Perusall and contribute substantive comments or replies to the discussion by noon Eastern Time on Thursday, September 17, 2026.

Thursday, September 24, 2026
Recursive Self-Improvement and AI Trajectories 🔗
Guest lecturers: Dwarkesh Patel and Daniel Kokotajlo
Thursday, October 1, 2026
Economic Impact of AI 🔗
Guest lecturers: Chad Jones and Erik Brynjolfsson
Thursday, October 8, 2026
Reinforcement Learning for Post-Training and Alignment 🔗
Guest lecturer: John Schulman
Thursday, October 15, 2026
Model Policies 🔗
Guest lecturer: Ziad Reslan
Thursday, October 22, 2026
Open-Source Models 🔗
Guest lecturer: Nathan Lambert
Thursday, October 29, 2026
TBD 🔗
Thursday, November 5, 2026
AI Interpretability 🔗
Guest lecturer: Jack Lindsey
Thursday, November 12, 2026
Alignment in the Age of Recursive Self-Improvement 🔗
Guest lecturer: Jakub Pachocki
Thursday, November 19, 2026
AI Biosecurity and Threat Modeling 🔗
Guest lecturer: Luca Righetti
No lecture on Thursday, November 26 — Thanksgiving Break
Thursday, December 3, 2026
TBD 🔗
Previous versions: Fall 2025 AI Safety Spring 2023 ML Theory Seminar Spring 2021 ML Theory Seminar