ZICQ
中 Log in / Sign up
Newsroom Research & Papers #AGI #Multimodular #Smart #World Model #A bitter lesson

AGI Is Not Multimodal: Brown Researcher Challenges Scaling Approach, Advocates Embodied Interaction

Avatar of Mr.Xu

By Mr.Xu

Published: · 6 views

中文阅读 (Chinese) English Version

Summary:In a thought-provoking article on The Gradient, Brown University PhD candidate Benjamin A. Spiegel argues that the current multimodal approach to AGI is doomed. He contends that LLMs learn syntactic heuristics rather than true world models, and that multimodal models sever deep connections between modalities while lacking embodied data. He advocates for embodied interaction as primary, allowing modality-specific processing to emerge naturally. The piece challenges scale-maximalist interpretation


Core Thesis: Multimodal Patching Cannot Lead to AGI

Benjamin A. Spiegel, a PhD candidate at Brown University, published a long-form article on The Gradient titled "AGI Is Not Multimodal," systematically arguing that the current approach of stitching together multimodal models to achieve AGI is fundamentally flawed. He contends that true AGI must be able to solve problems in the physical world, such as repairing a car, untying a knot, or preparing food, which require embodied intelligence grounded in a physical world model, not mere symbol manipulation.

LLMs Do Not Truly Understand the World

Spiegel points out that despite LLMs' impressive performance on language tasks, evidence suggests they do not learn true world models. He cites studies like OthelloGPT, showing that models may learn heuristic rules from training data rather than causal structures of the world. For example, OthelloGPT learned statistical regularities like "if B4 does not appear before A4, then B4 is empty," rather than the full state of an Othello board. He further argues that linguistic descriptions cannot fully infer the state of the physical world, so LLMs' "world models" are more like memorization of syntactic structures than understanding of semantics and pragmatics.

Three Major Flaws of Multimodal Approaches

Spiegel identifies three main problems with current multimodal methods:

  1. Deep connections between modalities are severed: During pretraining, each modality is optimized independently, and joint embedding spaces cannot capture complex one-to-many relationships between modalities, such as image captions at different abstraction levels or different physical action implementations of the same instruction.
  2. Modality partitioning does not match cognitive nature: In human cognition, reading, vision, language, and movement are supported by overlapping cognitive structures. Treating images and text as separate observation streams, and text generation and motion planning as separate action capabilities, hinders the discovery of more fundamental cognitive processes.
  3. Lack of concept innovation capability: Models merely replicate existing human conceptual structures rather than learning the ability to form new concepts from experience, limiting their generalization beyond the training distribution.

Reinterpreting the Bitter Lesson

Spiegel argues that Sutton's Bitter Lesson is often misinterpreted as opposing any structural assumptions, but in reality, structural assumptions like CNNs' translation invariance and Transformers' attention mechanism have been key to AI breakthroughs. He advocates for designing environments where modality-specific processing emerges naturally, rather than presupposing modality structures. He cites his own research on visual theory of mind, showing how abstract symbols emerge from communication among image-classifying agents.

Conclusion: From Mathematical to Conceptual Problem

Spiegel concludes that the hardest mathematical part of AGI (universal function approximators) is solved, and the remaining challenge is to inventory the needed functions and arrange them coherently—a conceptual, not mathematical, problem. He suggests either intentionally fusing modalities (drawing on human intuition and classical research) or redefining learning as an embodied interactive process where modalities naturally fuse. Though efficiency may decrease, flexible cognitive ability will be gained.

Industry Impact and Implications

This article provides important reflection for the AI community: scaling is not a panacea, and achieving AGI may require a fundamental paradigm shift. For developers, it suggests focusing on embodied intelligence, world models, and concept formation rather than blindly pursuing multimodal model scale.


Source: The Gradient, "AGI Is Not Multimodal", 2025. Link

— END —

Tags: #AGI #Multimodular #Smart #World Model #A bitter lesson

Community Comments

Loading live comments and annotations…