Gal Mishne - From Explanations to Mechanisms: Interpreting Computation in Graph Neural Networks

Gal Mishne - From Explanations to Mechanisms: Interpreting Computation in Graph Neural Networks

🎙 Gal Mishne 👥 42K 📅 September 3, 2026 ⏱ 52 min 👁 3 📄 original study 🧭 2026-09-03
Available in: English (current) Français

Keywords

Graph Neural NetworksInterpretabilityMechanistic InterpretabilityNeural Algorithmic ReasoningCircuit Discovery

Summary

Gal Mishne presents two complementary approaches to interpreting Graph Neural Networks (GNNs). The first approach focuses on post-hoc explanations, establishing theoretical connections between perturbation-based methods (like GNN Explainer) and edge gradients. Under a stability condition, the GNN Explainer objective reduces to selecting edges with positive gradients. Empirical results show high similarity between these methods, and layer-wise gradients are equivalent to occlusion for linear networks. This enables extracting path information from the computation graph, improving accuracy on a synthetic infection task. The second approach introduces MINAR (Mechanistic Interpretability for Neural Algorithmic Reasoning), which aims to identify internal circuits in GNNs trained on algorithmic tasks. By unrolling the network into a model computation graph and using activation patching, they extract circuits that are shared across nodes and tasks. They demonstrate that circuits form and are pruned during training, and that related tasks share circuit components. The talk highlights the challenges of adapting mechanistic interpretability to GNNs, such as handling shared parameters and graph-structured inputs.

162 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into GNN interpretability, bridging theoretical analysis with practical methods. The argumentation is solid, supported by theoretical results and empirical evaluations on standard datasets. The connection between perturbation and gradient methods is a significant contribution, simplifying explanation methods while maintaining accuracy. The MINAR framework extends mechanistic interpretability to GNNs, offering a novel perspective on understanding internal computations. The presentation is well-structured, clearly motivating each step and addressing potential limitations.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, presenting original research with theoretical proofs and empirical validation. The speaker references prior work and builds upon established methods. The title accurately reflects the content, which progresses from explanations to mechanisms. The talk is part of a workshop on interpretability, indicating relevance to the field. No external sources are cited beyond the workshop link, but the research is presented with sufficient detail for a conference talk.

159 words

Title / Content Match

The title accurately reflects the content, which transitions from post-hoc explanations to mechanistic interpretability of GNNs.

Quality & Reliability

8/10

The talk presents original research with theoretical results and empirical validation, delivered by an academic expert at a recognized workshop. The methods are clearly described, and the claims are supported by experiments on standard datasets. However, the presentation is a conference talk, so details are limited and some results are not fully peer-reviewed yet.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk contributes original theoretical results linking perturbation-based and gradient-based explanation methods for GNNs, showing that GNN Explainer’s objective reduces to positive edge gradients under a stability condition. It also introduces MINAR, a framework for mechanistic interpretability of GNNs, enabling circuit discovery in algorithmic reasoning tasks. This extends mechanistic interpretability beyond LLMs to graph-structured data.

Pour aller plus loin :

  • Mechanistic Interpretability — Overview of interpretability methods, including mechanistic approaches.
  • Graph Neural Networks — Background on GNN architectures and applications.
  • Neural Algorithmic Reasoning — Paper on using neural networks to execute algorithms, relevant to MINAR.
  • Activation Patching — Explanation of activation patching technique used in circuit discovery.

107 words

Radar Profile

The radar profile shows high scores in technical level and information quality, reflecting the advanced and rigorous nature of the content. The fiabilite_globale is also high, indicating strong trustworthiness. The quantite_information is slightly lower, as the talk is a conference presentation with limited time for exhaustive detail.

Reliability 8/10