PIRAMID
Physics-Informed Research for Ambitious Mechanistic Interpretability Development
We believe that the pursuit of ambitious mechanistic interpretability (AMI) supported by a rigorous science of AI systems is essential for scalable AI alignment. Starting with the methods and frameworks of statistical physics, our research will simultaneously build up the scientific foundations of three pillars of AI safety.
These are deeply interconnected: theory underpins applications, which provide empirical support for theoretical predictions, and principled validation methods facilitate quick feedback between the two.
PIRAMID is a participating group in the PIAMI coordinated research program. Our current work focuses on how neural networks learn and leverage hierarchical structure in real-world data, grounded in several hypotheses:
- Learned features are organized according to a hierarchical, scale-dependent notion of relevance.
- Faithful interpretability tools leverage this hierarchy.
- A renormalization-like framework can place principled, probabilistic bounds on cross-scale mechanistic behavior, potentially enabling worst-case guarantees.
We use these ideas a guiding heuristics, rather than strict prescriptions, and are open to using a wide variety of tools and methods.
Posts from PIRAMID
Our Team

Lauren Greenspan
Technical Director

Dmitry Vaintrob
Research Lead

Nischal Mainali
Research Affiliate

Ari Brill
Research Lead

Andrew Mack
Research Lead

Jennifer Lin
Research Affiliate