Schedule

Most materials for each day are posted in #intensive-october.

The day

Hour by hour

10:00
Lecture starts
12:30
Lunch at Mox
13:30
Afternoon session
17:40
Feedback form
18:00
Day ends

The buildings

Where

From Mon 5 Oct
Mox, 1680 Mission Street, San Francisco, CA 94103

Week 1 · 5–9 October

1 Mon 5 Oct Intro to ML Engineering · with transformer maths Dan Wilhelm Mox

Two tracks running in parallel: an introduction to deep learning for those who want it, and Claude Code for those who do not.

Prerequisites
  • A laptop, with the Claude Code CLI already working on it
  • gradient descent
  • linear algebra
2 Tue 6 Oct AI Alignment Introduction Margot Stakenborg Mox

Two combined challenges: choosing an alignment target, our vision for how systems should behave, and solving the technical problem of aligning systems with it. Frameworks for decomposing the technical problem (training stories, outer and inner alignment, inductive biases) and the critiques of them; how goal-directedness raises the stakes; and a survey of how hard and how severe people take the problem to be, alongside the spread of solution approaches the field pursues.

Prerequisites

Nothing technical beyond the foundations. The worldview reading in the prerequisites document.

Materials
3 Wed 7 Oct Mechanistic Interpretability Dan Wilhelm Mox

Methods for identifying features and circuits, the assumptions underneath them, and what has been found so far. Early CNN work, then transformers: linear probes, steering vectors and directional ablations; superposition and sparse autoencoders, how they are evaluated and where they fail; circuit discovery through logit attribution, the logit lens, path patching, ACDC, attribution graphs and causal scrubbing. Hands-on throughout, closing with a group discussion of the major critiques of the field.

Prerequisites

Nothing beyond the foundations.

Materials
4 Thu 8 Oct Alignment in Practice Margot Stakenborg Mox

Pretraining, post-training, deployment: what affordances each phase gives us for alignment, illustrated with the state-of-the-art methods at the major labs, plus the empirical results worth carrying around about how LLMs actually behave. The deployment section takes a higher-level view and covers the non-technical parts as well: responsible scaling policies, safety cases, the economic impact of AI, and governance.

Prerequisites

Nothing beyond the foundations.

Materials

Module notes(in progress as of August)

5 Fri 9 Oct Decision Theory and Reinforcement Learning David Quarel Mox

Reinforcement learning motivated from preferences: what the von Neumann-Morgenstern axioms assume, how preferences become a utility function, and how that becomes the reward, return and policy vocabulary of a Markov decision process.

Prerequisites
  • partial orders
  • expected utility

Week 2 · 12–16 October

6 Mon 12 Oct Steganography and Backdoors Julian Schulz Mox

Steganography as the concealment of hidden messages inside innocent-looking outputs, with cryptographically secure schemes that are computationally undetectable, and channels such as paraphrasing that destroy hidden communication, plus why steganographic behaviour is hard to detect or prevent, Merlin–Arthur classifiers as a possible counter-strategy, and the result that perfect steganography requires secret randomness satisfying H(M) ≤ H(K). Then cryptographic backdoors in LLMs: unelicitable backdoors resting on computational hardness, and white-box-undetectable approaches that hide triggers in random weight distributions.

Prerequisites
  • conditional entropy
  • lossy and lossless compression
  • the multivariate normal: moments and density
  • concentration inequalities: Markov and Chebyshev
  • pseudo-random functions
  • one-way functions
  • public-key encryption
Materials
7 Tue 13 Oct Mysteries of Deep Learning Charles Renshaw-Whitman and Brianna Grado-White Mox

A tour of the results that make deep learning strange. Each phenomenon is stated precisely, the experiment that shows it is described, and the candidate explanations are weighed against it. The point is not a catalogue: it is to leave you able to say what a theory of deep learning would have to account for, which is the standard the rest of the month is measured against.

Prerequisites
  • empirical risk minimisation
  • the bias-variance decomposition
  • overparameterisation
  • scaling laws
  • gradient descent

A short primer on Solomonoff induction and AIXI may open this day; where that material sits is still being decided.

8 Wed 14 Oct Training Dynamics Brianna Grado-White and Charles Renshaw-Whitman Mox

Implicit regularisation: how the training process itself biases toward simple solutions, via loss-landscape geometry, the edge of stability, simplicity bias, the neural tangent kernel, and the lazy versus rich regimes in deep linear networks as a toy model of deep learning. Emergence: grokking, induction heads and silent alignment, read through phase transitions (grokking as a transition from the lazy to the rich regime) and what that perspective buys us for detecting emergent capabilities early, which is the safety payoff.

Prerequisites
  • the singular value decomposition
  • ordinary differential equations
Materials
9 Thu 15 Oct Singular Learning Theory Tudor Dimofte Mox

Within an architecture, certain weight vectors correspond to structurally simpler networks. These degeneracies complicate the map from parameter space to function space, and make learning in neural networks substantially richer than learning in classical statistical models. Qualitative definitions of degeneracy through the parameter–function map, the Fisher information matrix and the curvature of the loss landscape; then the central quantitative definition, the local learning coefficient, from volume scaling asymptotics; then Watanabe’s free energy formula for Bayesian inference as a case study.

Prerequisites
  • Bayesian statistics: prior, posterior, likelihood
  • multivariate integrals and change of variables
Materials
10 Fri 16 Oct Data Attribution Brianna Grado-White and Charles Renshaw-Whitman Mox

The focus moves from weight space to training data, framed as the counterfactual impact of reweighting individual data points. Three frameworks each read the data-to-model map differently: influence functions, as an implicit function of data weights at a unique minimum; Bayesian influence functions, as a posterior distribution over parameters; and unrolling, as a concrete optimisation trajectory. They turn out to be closely connected (influence functions emerge as a limiting case of both alternatives) and the degeneracy phenomena from the SLT day reappear exactly where the classical theory breaks down.

Week 3 · 19–23 October

11 Mon 19 Oct Computational Mechanics Paul Riechers Mox

Starting from hidden Markov models, the module motivates generalised HMMs through minimality and uniqueness, then develops two views of Bayesian inference over emissions: belief states geometrically, and the mixed state presentation algorithmically. You design your own processes, look at the evidence that transformers trained on GHMM data learn belief-state geometry in their residual streams, and build mechanistic hypotheses about how that geometry gets constructed and used.

Prerequisites
  • Markov chains and row-stochastic matrices
  • hidden Markov models
  • the probability simplex
  • conditional probability and Bayes' rule
  • linear probes
Materials
12 Tue 20 Oct World Models Alec Boyd Mox

World models are what let an advanced agent plan and weigh counterfactuals without acting. Three units: a general introduction to world models and how they are used in modern AI; how they can be formalised in a reinforcement learning setting; and how agents use them to build abstractions. The treatment is deliberately interdisciplinary, combining computer science with principles from statistical physics, neuroscience and cognitive science.

Prerequisites
  • mutual information
  • causal graphs and interventions
Materials
13 Wed 21 Oct Project proposal day Leonard Bereska Mox

Pick a research area and write a proposal in it. The areas on the table by now: steganography and backdoors, mysteries of deep learning, training dynamics, singular learning theory, data attribution, and computational mechanics with world models.

14 Thu 22 Oct Physics for Deep Learning Tudor Dimofte Mox

Field-theoretic treatment of wide networks: perturbative expansion in the inverse width, the infinite-width Gaussian limit, the partition function correspondence, renormalisation and relevant operators, and why the depth-to-width ratio matters.

Prerequisites
  • statistical mechanics: the partition function and free energy
  • Gaussian integrals and the central limit theorem
  • second-order Taylor and perturbative expansion
  • Bayesian inference: prior, likelihood, posterior
  • kernels and Gram matrices
15 Fri 23 Oct SAEs and Activation Geometry Eric Michaud Mox

Sparse autoencoders are the most-used tool for turning activations into something a person can read, and also the most-argued-about. This day covers how they are trained, what a feature recovered by one is and is not evidence of, and the geometric picture underneath: superposition, the linear representation hypothesis, and where each of them is known to break. It builds on mechanistic interpretability in week 1 and feeds abstractions and latents in week 4.

Prerequisites
  • autoencoders
  • sparsity and L1 regularisation
  • residual stream
  • superposition
  • the linear representation hypothesis

Mechanistic interpretability (day 3) is the day this one assumes.

Week 4 · 26–30 October

16 Mon 26 Oct Abstractions and Latents Satya Benson Mox

Formal theory of natural latents (mediation and redundancy) and the condensation framework, with their agreement and translatability theorems.

Prerequisites
  • Bayesian networks
17 Tue 27 Oct Agent Foundations Aden Power Mox

A central difficulty in alignment is that we have to reason about the behaviour of systems that do not exist yet, and cannot learn from our mistakes with them. This module develops formal tools for doing so across several agendas: coherence arguments and the complete class theorem, Löb's theorem and the Löbian obstacle to safe self-modification, tiling agents and Vingean reflection, logical induction and reasoning under logical uncertainty, functional and updateless decision theory, and the thermodynamics of optimisation.

Prerequisites
  • basic formal logic: provability and quantifiers
  • basic computability: programs and halting
  • elementary discrete probability
Materials
18 Wed 28 Oct To be confirmed Mox

To be confirmed.

19 Thu 29 Oct To be confirmed Mox

To be confirmed.

20 Fri 30 Oct Project proposal day Leonard Bereska Mox

The second proposal day. The areas added since the first: physics of deep learning, SAEs and activation geometry, abstractions and latents, agent foundations, and debate.

Prerequisites

Fundamentals

from prerequisites document

Linear algebra
vectors and matrices, rank, null spaces, the rank-nullity theorem, orthogonality, invertibility, positive definiteness, eigenvalues, spectral decomposition, the singular value decomposition
Calculus
limits, derivatives and integrals, partial and directional derivatives, gradients, Jacobians, the chain rule in several variables, the Hessian, second-order Taylor expansion, multivariate integrals, change of variables, O and o notation
Probability
joint and conditional probability, Bayes' rule, the probability simplex, expectation, variance, moments, independence, the law of large numbers, the multivariate normal
Information theory
entropy, mutual information, KL divergence, cross-entropy
Deep learning
loss functions (cross-entropy and squared error), backpropagation, stochastic gradient descent, ReLU and softmax, multi-layer perceptrons, the inputs and outputs of a transformer, weights and activations, training, validation and test sets, hyperparameters, optimisers, overfitting and underfitting