Mahdi Soltanolkotabi - Interpreting Generative Models for Better Steering: Generation to Verifiable

Mahdi Soltanolkotabi - Interpreting Generative Models for Better Steering: Generation to Verifiable

🎙 Mahdi Soltanolkotabi 👥 42K 📅 September 3, 2026 ⏱ 42 min 👁 2 📄 original study 🧭 2026-09-03
Available in: English (current) Français

Keywords

diffusion modelssparse autoencodersinterpretabilityvisual reasoningcounting

Summary

Mahdi Soltanolkotabi presents research on interpreting and steering diffusion models to improve visual reasoning, particularly object counting. The talk begins by highlighting persistent failures in visual reasoning, even in advanced models. The first part focuses on training sparse autoencoders (SAEs) on the hidden representations of Stable Diffusion v4 to disentangle features. Key findings include that image layout emerges very early in the denoising process, while style emerges later. This insight enables early intervention. The second part introduces ‘Aphina’, a steering method that uses an object detector at intermediate diffusion steps to detect counting errors and then modifies the noise prediction using a control prompt (e.g., removing the number from the prompt) to correct the count. The method is computationally efficient as it intervenes early, unlike proprietary models that appear to generate the full image and then correct it. Experiments on a custom dataset with 360 prompts and four complexity levels show significant accuracy improvements over base models with minimal added compute. The talk concludes by mentioning ongoing work on applying RLVR (Reinforcement Learning with Verifiable Rewards) to broader visual reasoning tasks.

181 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the internal dynamics of diffusion models, showing that SAEs can reveal when different semantic features (layout vs. style) emerge. This is a significant contribution to interpretability research, which has largely focused on language models. The proposed steering method, Aphina, is a practical application of these insights, offering a compute-efficient alternative to existing approaches. The argumentation is solid: the speaker motivates the problem with concrete examples, explains the methodology clearly, and supports claims with experimental results. The comparison with proprietary models’ reasoning traces is speculative but well-flagged. The main limitation is that the talk presents work in progress, and some technical details are glossed over.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, presenting original research with a clear methodology. The speaker references standard interpretability tools (SAEs, J-Lens from Anthropic) and existing literature on counting issues. The quality of sources is high, as the work is presented at an IPAM workshop, a reputable institution. The title accurately reflects the content, covering both interpretation and steering. The talk does not cite specific papers, but the description links to the workshop page, which likely contains further references. The speaker’s claims are appropriately hedged, and the experimental setup is described in sufficient detail to assess validity.

220 words

Title / Content Match

The title accurately reflects the content: the talk covers interpreting generative models (diffusion models) via sparse autoencoders and using these insights for steering, including a method for verifiable reasoning (counting).

Quality & Reliability

8/10

The talk presents original research with a clear methodology, including training sparse autoencoders on diffusion models and proposing a novel steering mechanism. The claims are supported by experimental results on custom and existing datasets. However, the talk is a presentation of ongoing work, and some details are omitted, limiting full reproducibility. The speaker is a recognized researcher, and the venue (IPAM) is reputable.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk contributes original research on interpreting diffusion models using sparse autoencoders, revealing that layout emerges early in the denoising process. This insight leads to a novel, compute-efficient steering method (Aphina) for correcting object counting errors, which is a significant improvement over existing complex pipelines. The work also highlights the inefficiency of current proprietary models, which appear to generate full images before correcting errors.

Pour aller plus loin :

117 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a technically deep, reliable, and information-rich presentation. The balance between quantity and quality of information is strong, with a slight emphasis on technical level, reflecting the advanced nature of the content.

Reliability 8/10