2026

  1. Elicitation without Backpropagation: Steering Model Behavior by Optimizing the Latent Posterior

    Garrett Baker, Vinayak Pathak, Daniel Murfet, Susan Wei

    · Preprint · arXiv

  2. Interpreting Reinforcement Learning Agents with Susceptibilities

    Chris Elliott*, Einar Urdshals*, David Quarel, Daniel Murfet

    · Preprint · arXiv

  3. Linear Response Estimators for Singular Statistical Models

    Chris Elliott, Daniel Murfet

    · Preprint · arXiv

  4. Susceptibilities and Patterning: A Primer on Linear Response in Bayesian Learning

    Chris Elliott, Daniel Murfet

    · Preprint · arXiv

  5. Spectroscopy at Scale: Finding Interpretable Structure in Pythia-1.4B

    Daniel Murfet, Andrew Gordon, Max Adam, George Wang, Jesse Hoogland, Garrett Baker, William Snell, Stan van Wingerden, Adam Newgas, Billy Snikkers, Rohan Hitchcock

    · Technical report · Timaeus

  6. How to Scale Susceptibilities

    Andrew Gordon, Max Adam, Jesse Hoogland, Daniel Murfet

    · Technical report · Timaeus

  7. Interpreting the Ising Model

    Andrew Gordon, Rohan Hitchcock, Daniel Murfet

    · Technical report · Timaeus

  8. Guide for Sampling Hyperparameter Selection

    Rohan Hitchcock, Jesse Hoogland, George Wang, Andrew Gordon

    · Technical report · Timaeus

  9. Patterning: The Dual of Interpretability

    George Wang, Daniel Murfet

    · Preprint · arXiv

  10. Towards Spectroscopy: Susceptibility Clusters in Language Models

    Andrew Gordon*, Garrett Baker*, George Wang, William Snell, Stan van Wingerden, Daniel Murfet

    · Preprint · arXiv

  11. Stagewise Reinforcement Learning and the Geometry of the Regret Landscape

    Chris Elliott, Einar Urdshals, David Quarel, Matthew Farrugia-Roberts, Daniel Murfet

    · Preprint · arXiv

2025

  1. Compressibility Measures Complexity: Minimum Description Length Meets Singular Learning Theory

    Einar Urdshals*, Edmund Lau*, Jesse Hoogland, Stan van Wingerden, Daniel Murfet

    · Preprint · arXiv

  2. Influence Dynamics and Stagewise Data Attribution

    Jin Hwa Lee*, Matthew Smith*, Maxwell Adam, Jesse Hoogland

    · Preprint · arXiv

  3. The Loss Kernel: A Geometric Probe for Deep Learning Interpretability

    Maxwell Adam*, Zach Furman*, Jesse Hoogland

    · Preprint · arXiv

  4. Bayesian Influence Functions for Hessian-Free Data Attribution

    Philipp Alexander Kreer, Wilson Wu, Maxwell Adam, Zach Furman, Jesse Hoogland

    · Preprint · arXiv

  5. Embryology of a Language Model

    George Wang, Garrett Baker, Andrew Gordon, Daniel Murfet

    · Preprint · arXiv

  6. From Global to Local: A Scalable Benchmark for Local Posterior Sampling

    Rohan Hitchcock, Jesse Hoogland

    · Preprint · arXiv

  7. Structural Inference: Interpreting Small Language Models with Susceptibilities

    Garrett Baker*, George Wang*, Jesse Hoogland, Vinayak Pathak, Daniel Murfet

    · Preprint · arXiv

  8. Modes of Sequence Models and Learning Coefficients

    Zhongtian Chen, Daniel Murfet

    · Preprint · arXiv

  9. Programs as Singularities

    Daniel Murfet, Will Troiani

    · Preprint · arXiv

  10. You Are What You Eat – AI Alignment Requires Understanding How Data Shapes Structure and Generalisation

    Simon Pepin Lehalleur*, Jesse Hoogland*, Matthew Farrugia-Roberts*, Susan Wei, Alexander Gietelink Oldenziel, George Wang, Stan van Wingerden, Zach Furman, Liam Carroll, Daniel Murfet

    · Preprint · arXiv

  11. Structure Development in List-Sorting Transformers

    Einar Urdshals, Jasmina Urdshals

    · ICML 2025 SMUNN Workshop · arXiv

  12. Dynamics of Transient Structure in In-Context Linear Regression Transformers

    Liam Carroll, Jesse Hoogland, Matthew Farrugia-Roberts, Daniel Murfet

    · Preprint · arXiv

2024

  1. Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient

    George Wang*, Jesse Hoogland*, Stan van Wingerden*, Zach Furman, Daniel Murfet

    · ICLR 2025 · Spotlight · ICLR 2025 · arXiv

  2. Loss Landscape Degeneracy and Stagewise Development of Transformers

    Jesse Hoogland*, George Wang*, Matthew Farrugia-Roberts, Liam Carroll, Susan Wei, Daniel Murfet

    · TMLR · Best Paper at 2024 ICML HiLD Workshop · TMLR · arXiv

2023

  1. The Local Learning Coefficient: A Singularity-Aware Complexity Measure

    Edmund Lau*, Zach Furman*, George Wang, Daniel Murfet, Susan Wei

    · AISTATS 2025 · AISTATS 2025 · arXiv

* Equal contribution.