Research
26 publications, 2023–2026. See also our events.
2026
-
Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior
-
Interpreting Reinforcement Learning Agents with Susceptibilities
-
Linear Response Estimators for Singular Statistical Models
-
Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning
-
Spectroscopy at Scale: Finding Interpretable Structure in Pythia-1.4B
-
How to Scale Susceptibilities
-
Interpreting the Ising Model
-
Guide for Sampling Hyperparameter Selection
-
Patterning: The Dual of Interpretability
-
Towards Spectroscopy: Susceptibility Clusters in Language Models
-
Stagewise Reinforcement Learning and the Geometry of the Regret Landscape
2025
-
Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory
-
Influence Dynamics and Stagewise Data Attribution
-
The Loss Kernel: A Geometric Probe for Deep Learning Interpretability
-
Bayesian Influence Functions for Hessian-Free Data Attribution
-
Embryology of a Language Model
-
From Global to Local: A Scalable Benchmark for Local Posterior Sampling
-
Structural Inference: Interpreting Small Language Models with Susceptibilities
-
Modes of Sequence Models and Learning Coefficients
-
Programs as Singularities
-
You Are What You Eat – AI Alignment Requires Understanding How Data Shapes Structure and Generalisation
-
Structure Development in List-Sorting Transformers
-
Dynamics of Transient Structure in In-Context Linear Regression Transformers
2024
-
Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient
-
Loss Landscape Degeneracy and Stagewise Development of Transformers
2023
-
The Local Learning Coefficient: A Singularity-Aware Complexity Measure
* Equal contribution.