Core Event: LeCun Formalizes JEPA as the World Model Core Architecture

At the European Conference on Computer Vision 2026 (ECCV 2026), Yann LeCun, Turing Award laureate and one of the three founding figures of deep learning, delivered the keynote speech titled World Models: Enabling the next AI revolution. In it, he systematically laid out ten years of thinking on world models and the trajectory toward true artificial general intelligence (AGI). LeCun explicitly rejected the prevailing generative approach to world modeling, classifying pixel-level prediction as an impossible task, and reaffirmed Joint Embedding Predictive Architecture (JEPA) as the viable path forward. He also announced the formation of Advanced Machine Intelligence (AMI Labs), a company aimed at accelerating engineering progress on this technical route.
Challenging the Mainstream: Why “Predicting Pixels” Is Fundamentally Flawed

LeCun argued that many teams currently attempting to build world models using video generation models (e.g., Sora, Genie) are fundamentally mistaken. He pointed out that pixel-level future prediction is inherently unsolvable: the current frame captured by a camera does not contain all information needed to determine the next frame—such as whether someone enters outside the camera’s view or how lighting changes—making true prediction impossible. Even techniques like VAE or latent variables only mitigate blurring temporarily; the model still “imagines” nonexistent futures and fails to grasp the underlying physics.
He illustrated this with autonomous driving: training a system to predict the next camera frame forces it to learn irrelevant details—quantum-level foliage movements, water ripples, or clothing wrinkles—information irrelevant to decision-making. This wastes model capacity and training data, degrading the system’s representation capability.
JEPA Architecture: Making Plan-Based Predictions in Abstract Space
JEPA’s core philosophy is “discard what cannot be predicted.” Its workflow first maps raw inputs (e.g., image frames, action commands) to a compact representation space via an encoder, then performs prediction directly in that space. This transformation inherently filters out high-dimensional noise and unobservable variables, transforming an intractable prediction problem into a solvable one.
The architecture naturally supports hierarchical planning. LeCun emphasized that humans do not plan journeys millisecond-by-millisecond (e.g., muscle contractions), but across multiple abstraction levels: high-level goals (“arrive at the airport”), mid-level subgoals (“take a taxi”), and low-level routines (“stand up”). JEPA’s representation space enables such hierarchical modeling: high-level representations preserve core plan-relevant structure, while low-level ones handle execution details.
Key Contrast: Text Training vs. Vision-Physical Experience

LeCun demonstrated the fundamental limitation of current paradigms through information volume:
- Typical large language model pretraining: ~30 trillion tokens (~10¹⁴ bytes), equivalent to 400,000 years of human reading
- A four-year-old child’s visual-tactile input: ~10¹⁴ bytes (2 million retinal nerve fibers × 1 byte/sec × 60,000 hours awake)
This means text-only models have an intrinsic learning ceiling far below that of human infants. He underscored that computer vision researchers especially understand the physical world’s “dirtiness,” continuity, and high dimensionality—explaining why large language models fail catastrophically when deployed in reality.
Practical Guidance

- For early explorers: Research teams working on world models, embodied intelligence, or robot decision planning should closely examine JEPA’s implementation in representation invariance and hierarchical optimization
- For those waiting: Expectations for end-to-end autonomous driving or household robotics should remain realistic—LeCun himself acknowledges L5 autonomy remains unsolved, and the industry still relies on hybrid solutions with explicit physics engines
Final Thoughts
LeCun’s speech represents a fundamental recalibration of AI development trajectories. By sharply distinguishing between “fluent” declarative knowledge (trained responses) and “推理” inferential ability (solving novel problems ad hoc), he positions inference as the essence of intelligence. If the JEPA path is validated, agents over the next few years will evolve from “pixel mimics” to “physical law reasoners,” potentially reshaping the entire paradigm of embodied AI research.
