Ryan Cotterell - Two views of a language model - IPAM at UCLA

Ryan Cotterell - Two views of a language model - IPAM at UCLA

🎙 Ryan Cotterell 👥 42K 📅 September 1, 2026 ⏱ 45 min 👁 22 📄 literature review 🧭 2026-09-01
Available in: English (current) Français

Keywords

language modelcomputationexpressivitysuccinctnessUHAT

Summary

Ryan Cotterell presents two complementary views of language models: statistical (probability distribution over strings) and computational (model of computation). He argues that these views diverge on rare or structured inputs, where average-case prediction does not guarantee algorithmic correctness. He introduces formal definitions: a language model as a probability measure over strings, and the autoregressive factorization as a theorem, not a definition. He then discusses the computational view, proposing modes of acceptance to convert distributions into recognizers. He emphasizes the need for mathematical idealizations of transformers, such as fixed-precision and hard attention, to enable formal analysis. He introduces the unique-hard-attention transformer (UHAT) and states key results: fixed-precision UHATs recognize exactly the star-free languages, placing them below RNNs in expressivity. He argues that succinctness is a sharper measure than raw expressivity, showing that UHATs can be exponentially more succinct than LTL and RNNs, and doubly exponentially more succinct than finite automata. He also shows that verification problems for transformers are EXPSPACE-complete, implying no efficient general reasoning is possible. The talk concludes by highlighting the importance of choosing the right mathematical tools for analyzing transformers.

182 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights by clearly distinguishing statistical and computational perspectives on language models, a distinction often blurred in the field. It argues convincingly that average-case performance does not imply algorithmic correctness, using the sorting example to illustrate this point. The presentation of formal results on UHAT expressivity and succinctness is rigorous and well-motivated, offering a fresh lens on transformer capabilities. The argumentation is solid, building from definitions to theorems, and the emphasis on precision and idealization choices is well-justified. The talk successfully challenges the common narrative that transformers are computationally powerful, instead showing their limitations and the importance of succinctness.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, presenting formal definitions and theorems with clear statements. It references prior work, such as the Turing completeness results by Pérez et al., and situates its contributions within the FLAN community. The sources cited are appropriate and credible, though the talk does not provide full citations for all claims. The title accurately reflects the content, and the talk stays on topic throughout. The presentation is well-structured, with a clear progression from definitions to results. The only minor weakness is the lack of detailed proofs, which is expected in a workshop talk.

212 words

Title / Content Match

The title accurately reflects the content: the talk contrasts statistical and computational views of language models.

Quality & Reliability

8/10

Talk by a recognized researcher (ETH Zurich) at a prestigious workshop (IPAM). Presents formal results with mathematical rigor, but as a conference talk, it lacks full proofs and peer-reviewed publication details.

Key Moments

Cited Sources

Concurring Sources

  • Attention is Turing Complete — Referenced as a prior result on Turing completeness of attention, which the talk critiques due to unrealistic precision assumptions.

Dissenting Sources

  • Attention is Turing Complete — The talk argues that Turing completeness results rely on arbitrary precision, which is unrealistic for actual transformers, thus presenting a contrasting view.

Contribution & Novelties

The talk offers a novel synthesis of statistical and computational perspectives on language models, arguing for the primacy of the computational view for understanding algorithmic behavior. It presents recent results on UHAT expressivity and succinctness, showing that transformers are subregular but can be exponentially more succinct than other formalisms. This reframes the debate on transformer capabilities, suggesting that succinctness, not raw expressivity, is the key metric. The talk also highlights the importance of precision and idealization choices in theoretical analysis.

Pour aller plus loin :

126 words

Radar Profile

The radar profile shows high scores in information quality and technical level, reflecting the talk's rigorous formal content. The quantity of information is also high, but the reliability score is slightly lower due to the nature of a workshop talk without full peer review. Overall, the profile indicates a technically dense and informative presentation.

Reliability 8/10