Paul Riechers - The shape of beliefs and abstraction in neural networks - IPAM at UCLA

Paul Riechers - The shape of beliefs and abstraction in neural networks - IPAM at UCLA

🎙 Paul Riechers (Simplex) 👥 42K 📅 September 4, 2026 ⏱ 54 min 👁 16 📄 original study 🧭 2026-09-04
Available in: English (current) Français

Keywords

belief updatingworld modelfractal geometryabstractionmechanistic interpretability

Summary

Paul Riechers, co-founder of Simplex, presents research on the interpretability of neural networks, focusing on how next-token prediction leads to predictable low-dimensional fractal structures associated with beliefs and abstraction. He argues that transformers learn to perform Bayesian updates over latent states of a world model during inference, going beyond classical computational paradigms. The talk outlines five key findings: (i) models effectively perform Bayesian updates, (ii) the world model exceeds classical computation, (iii) networks factor the world into conditionally independent parts represented in orthogonal subspaces, (iv) multiple ergodic components induce sparsity, and (v) shared parts among components induce abstraction. Riechers demonstrates these insights using synthetic hidden Markov models, showing that the residual stream of transformers embeds belief states in a low-dimensional geometry, which can be extracted via PCA. He also discusses the architectural constraints that lead to fractal structures like the Sierpinski gasket in intermediate representations. The talk concludes with a call to action for building deep understanding of intelligence, emphasizing the urgency due to rapid AI capabilities and the need for interpretability foundations.

173 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides high-value insights by bridging theoretical predictions with empirical observations in neural networks. The argumentation is solid: the speaker starts with a clear falsifiable hypothesis, uses carefully controlled synthetic data to test it, and progressively refines the model. The use of hidden Markov models allows for exact calculations, and the demonstration that transformers learn to represent belief states in a low-dimensional subspace is compelling. The speaker addresses potential objections, such as whether the model is merely doing lookup, by emphasizing the representational structure and its implications for interventions and out-of-distribution behavior. The argument is strengthened by mechanistic analysis of a single-layer transformer, showing how architectural constraints lead to fractal geometries. The call to action for interpretability is well-motivated but somewhat philosophical, which may dilute the scientific focus.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is high: the research is presented as original work with peer-reviewed publications (NeurIPS 2024, ICML 2025). The methodology is transparent, and the claims are falsifiable. The speaker references prior work, such as Josh Batson’s talk, and builds on established concepts like Bayesian inference and hidden Markov models. The title accurately reflects the content, focusing on the shape of beliefs and abstraction. The talk is part of an IPAM workshop, which adds credibility. No external sources are cited beyond the workshop page, but the internal references to papers are sufficient for the context. The audience questions show engagement and the speaker responds with clarifications, indicating a rigorous discussion.

254 words

Title / Content Match

The title accurately reflects the content, which focuses on the geometric structure of beliefs and abstraction in neural networks.

Quality & Reliability

8/10

The talk presents original research with a clear methodology, falsifiable claims, and references to peer-reviewed publications (NeurIPS 2024, ICML 2025). The speaker demonstrates scientific rigor through explicit hypotheses, controlled experiments, and mechanistic analysis. While the claims are ambitious and not yet fully validated in large-scale models, the approach is transparent and grounded in established theory.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk presents a novel framework for understanding neural network representations as belief states over a world model, with a geometric structure that is fractal and low-dimensional. This goes beyond existing mechanistic interpretability work by providing a predictive theory that can be tested and used for interventions. The identification of architectural constraints leading to fractal geometries is a new insight. The talk also connects these findings to broader questions of abstraction and sparsity, offering a unified perspective.

Pour aller plus loin :

119 words

Radar Profile

The radar profile shows high scores in information quality, technical level, and reliability, with slightly lower scores in information quantity and overall reliability. This indicates a technically dense and rigorous presentation, but with a narrow focus that may limit its breadth.

Reliability 8/10

💬 No comments were provided for analysis.