ImplicCapital ← All Research
Implic Focus · March 2026

World Model: Another Path to AGI

If LLMs climb the south slope of AGI by understanding the world through language, World Models take the north slope — understanding the world through physical interaction.

Reading time ~15 min
Illustration of a hand-drawn globe with the title World Model: Another Path to AGI
Implic Notes

If LLMs are climbing the south slope of AGI by "understanding the world through language," then World Models take the north slope — understanding the world through physical interaction. The south slope is crowded with capital and converging on a paradigm; the north slope is relatively undervalued, but it's a bet on AI systems building internal representations of the physical world, predicting the consequences of actions, and ultimately achieving generalized robot control.

2024 was the funding inflection point for the World Model track: VC investment jumped from $6–8B in 2023 to $14–16B, surpassing $20B in 2025. The capital is in place, but the paradigm (data mix / algorithmic architecture) is far from converged — this is a game that appears to have many players, but where only a few can win, with a deployment timeline that may be longer than LLMs.

Key Takeaways:

01

Track & Funding: The Inflection Point Has Arrived, Concentration Is Extreme


1.1 What Is a World Model?

Any technical direction that attempts to have AI systems build an internal representation of the physical world, predict the consequences of actions, and thereby achieve generalized robot control can be classified under the World Model umbrella. It has two interpretations — narrow and broad:

Narrow vs. broad interpretations of "World Model"
Cells marked … are cut off in the source sheet
CategoryRepresentative DirectionsCore Idea
Narrow World ModelJEPA (LeCun), video prediction modelsPredict future states in latent or p…
Spatial IntelligenceWorld Labs (Fei-Fei Li)Reconstruct 3D world from 2D ima… relationships
Broad Embodied AIVLA, Generalist Policy, Hierarchical SystemsAll foundation-model approaches … perception to action
Exhibit 1Source: Implic Capital research

1.2 Funding Landscape: More Crowded Than LLM in 2015

Estimated VC investment in the World Model / embodied-AI track
US$B per year · solid bar = lower bound of estimate, lighter cap = estimate range
0 5 10 15 20 ~$4–5B ~$6–8B ~$10–15B $20B+ 2022 2023 2024 2025
Exhibit 2Source: PitchBook, Crunchbase, public reporting, Implic Capital research

Inflection points (2023H2 → 2024H1):

The driving force is capital overflow from the LLM track — robotics is one of the largest capital-capacity tracks after large language models. But when exactly the inflection point will arrive is a matter of extreme disagreement within the industry: Pi CEO Sergey Levine says at least ten years before seeing large-scale deployment; Sunday CEO Tony Zhao says 18 months; some World Model company CEOs are as optimistic as 8 months.

1.3 Key Players and Segmentation

Companies in the space can be understood along two dimensions: which use case they enter from (robot manipulation vs. autonomous driving vs. general 3D), and which layer of the technology stack they operate in (world understanding layer vs. decision execution layer vs. hardware layer).

A notable trend: value is migrating from the hardware layer to the model layer. Figure has the highest valuation ($39B), but its core narrative has already shifted from "making robots" to "robot operating system"; Skild and Generalist make no hardware whatsoever and purely sell models/software. This mirrors the PC era, where hardware companies (Dell) gave way to operating systems (Microsoft).

Value is migrating from the hardware layer to the model layer — a replay of the PC era, where Dell gave way to Microsoft.

A. General-Purpose World Models

General-purpose World Model startups
Cells marked … are cut off in the source sheet
CompanyValuationFundingApproach
AMI Labs$3.5B$1.03B SeedJEPA
World Labs$5B+$1.23B (incl. Feb 2026 $1B raise)Spatial Intelligence / 3D
OdysseyUndisclosed$27MGeneral World Model…
Decart$3.1B$153MReal-time generative … world model
Exhibit 3Source: PitchBook, Crunchbase, Implic Capital research

B. Robot Foundation Models

Robot Foundation Model startups
A further commentary column in the source sheet is cut off and omitted here
CompanyValuationFundingApproach
Skild AI$14B$1.83BGeneralist Policy
Physical Intelligence (Pi)$5.6B → in talks at $11B+~$1.1B+VLA + Flow Matching
Sunday Robotics$1.15B~$200MDiffusion Policy + crowdsourcing
Generalist~$3B$140MUniversal robot control
Exhibit 4Source: PitchBook, Crunchbase, Implic Capital research

C. Big Tech Landscape

Big tech is the most important variable in the space — theoretically best positioned, but with very different objectives:

D. Integrated Hardware-Software / Humanoid Robots

The source sheet attached to this segment duplicates the end-to-end vs. modular architecture comparison rebuilt in Section II (Exhibit 5); representative integrated hardware-software players named across this report include Figure, 1X, Tesla and Agility Robotics.

02

Technology Roadmap: More Divergence Than Consensus


Conclusion: Technical roadmaps have not converged, and will not converge in the near term. Three core points of divergence:

2.1 Training Framework: Pre / Mid / Post-train

Robot model training is divided into three stages, each with different data requirements:

2.2 System Architecture: End-to-End vs. Modular

End-to-end vs. modular (hierarchical) architectures
A status/assessment column in the source sheet is cut off and omitted here
ArchitectureTechnical LogicRepresentative Companies
End-to-EndA single model from pixels to actions; lets the model learn its own intermediate representations, avoiding human-designed information bottlenecksPi (pi0, 3B params), Google Gemini Robotics, Sunday
Modular (Hierarchical)Large upper-layer model for planning (0.1–2 Hz); small lower-layer model for execution (100–1000 Hz)Figure, Skild, 1X, Agility
Exhibit 5Source: Implic Capital research

Generalist Policy is the most aggressive sub-direction of the modular approach: training a single, cross-task, cross-embodiment general execution model (Skild is the representative). If it succeeds, it will play a role similar to ARM in the chip space. The risk: if the upper layer uses a general LLM and the lower layer uses scenario-specific small models, anyone can build this stack — the moat depends on data accumulation and depth of engineering integration.

2.3 Action Generation Methods: From Imitation to Generation

Whether end-to-end or modular, the model ultimately needs to output joint actions. This "final step" has gone through three generations of evolution in three years:

2.4 Four Paths for World Models

Technical paths for World Models
The source sheet is truncated at the right and bottom; cells marked … or n/a are cut off in the original
PathCore IdeaRepresentative
JEPAInstead of predicting pixels, predict abstract representations of future states in latent space. Focuses only on causal structure, ignores irrelevant detailsAMI Labs (LeCun + Xie Saining, $3.5B)
Video GenerationDirectly generate future video frames — essentially using video generation as a physics simulatorNVIDIA Cosmos (900 trillion token training), Google Genie 3, Decart
3D SpatialInstead of predicting the future, reconstruct a persistent 3D scene …n/a
Exhibit 6Source: Implic Capital research

JEPA / Video Gen focuses on the temporal dimension ("what happens next"); 3D Spatial Intelligence focuses on the spatial dimension ("what does the world look like"). The two are naturally complementary: first use 3D reconstruction to understand scene structure, then use temporal prediction models to plan actions.

2.5 World Model Companies and Embodiment

Hardware coupling models among World Model companies
ModeRepresentativeLogic
Tightly Coupled (Hardware-Software Integrated)Figure, 1X, TeslaEnd-to-end requires deep software-hardware coupling; own hardware = own data
Pure Software (Hardware-Agnostic)Pi, SkildBuild universal models/data to serve all hardware vendors; asset-light, high leverage
Exhibit 7Source: Implic Capital research

This is fundamentally determined by technical approach: companies pursuing end-to-end VLA are theoretically able to be pure software; those pursuing Hierarchical with lower-layer Policy tightly coupled to hardware are not. Notably, pure-software companies are also beginning to build their own hardware — primarily as data collection vehicles (e.g., Sunday's home robot collects data in real user environments through product sales) — but their core DNA remains non-hardware.

03

Data: A "Good Vantage Point" Before Technical Convergence


When paradigms have yet to converge, data is the most valuable lens for observing this space — because regardless of which technical path wins, data requirements will continue to grow as new use cases unlock. This is a structural demand, not a one-time event.

3.1 Three-Stage Data Requirements

Data requirements across the three training stages
A bottleneck/assessment column in the source sheet is cut off and omitted here
StageCore DataVolume & Cost
Pre-trainInternet video, games, simulationLarge volume; near-zero cost
Mid-trainOpen X-Embodiment (21 institutions, 22 robot types, 527 skills), UMI human motionModerate
Post-trainTeleoperation data$50–200/hour; tightly coupled to hardware
Exhibit 8Source: Implic Capital research

Core tension: Pre-train data volume is sufficient, but quality (especially action alignment) is not. Sergey Levine estimates it will take 10 years to see a "GPT moment" level of emergence.

Simulation data effectiveness decreases across stages: high for Pre-train, moderate for Mid-train, essentially zero for Post-train — unless simulation physics accuracy sees an order-of-magnitude improvement, which will not happen in the near term. This gap has also spawned a unicorn: Lightwheel, tightly integrated with NVIDIA.

3.2 The Dilemma of Data Vendor Business Models

Data companies face a structural contradiction: low standardization (every client has different hardware and data formats) but limited client budgets and long decision chains. This is not a business that can scale quickly like SaaS — it is more like high-end consulting.

Representative vendors by category (international + Chinese):

Teleoperation

First-Person Perspective

Tactile / Motion Capture

Data Platforms / Simulation

Track assessment: The real opportunity lies in end-to-end solutions (collection → labeling → training → iteration, full pipeline). Raw data volume alone is not a moat; data quality, scenario coverage, and model adaptation capability are the true differentiators.

04

Commercialization: B2B Barely Works, B2C at Least 5 Years Away


4.1 B2B: Currently Only Warehouse Logistics Is Scaling

Core assessment: Warehouse logistics is the scenario that is currently showing initial traction, but the essence is "replacing existing automation with a more expensive solution" — ROI has yet to be proven.

The most representative case is Agility Robotics + Amazon: Digit was tested at Amazon's Sumner warehouse for 18 months, achieving a 98% task success rate and operating costs of $10–12/hour (vs. $30/hour for human labor). It has been in full-time deployment at GXO Flowery Branch for a year, cumulatively moving over 100,000 totes. Agility completed a $400M Series C in 2025, with Salem factory production capacity targeting 10,000 units/year.

98%
Digit task success rate at Amazon Sumner
$10–12/hr
Operating cost vs. $30/hr human labor
100K+
Totes moved at GXO Flowery Branch
10K/yr
Salem factory capacity target

Automotive manufacturing is ramping up: BMW + Figure (Spartanburg factory), Mercedes-Benz + Apptronik (Apollo factory testing).

Key bottlenecks: Technically, generalization is insufficient ("passing a demo doesn't mean it works on the production line"); on the sales side, decision chains are long (robots are capex decisions; large-client sales cycles are 6–12 months); budget-wise, potential clients have limited budgets, while large clients have enough but impose strict requirements.

4.2 B2C: Won't Truly Open Before 2030

All current "home robot" narratives are fundraising stories. Key unlock conditions:

Unlock conditions for the home (B2C) scenario
"Expected timeline" column transcribed from a second capture of the same source sheet; its final cell is cut off (n/a)
ConditionCurrent StatusExpected UnlockExpected Timeline
Safety certification frameworkAlmost nonexistent2028–20302024–2026
Cost < $10KCurrently $50K+ (high-end) / $16K (low-end)2030+2025–2028
Operational reliability > 99.9%~80–90% (lab setting)2030+2027–2030
Scenario generalizationLimited to controlled environments2028–2030 (limited scenarios)n/a
Exhibit 9Source: Implic Capital research

The most likely path: industrial scenarios (2024–2028) → commercial service scenarios (hotels, hospitals, 2028–2030) → home scenarios (2030+).

05

Investment Perspective: A Game for the Few


Returning to the "north slope vs. south slope" metaphor at the opening — LLMs travel the road of "descriptions about the world"; World Models travel the road of "the world itself." Robots need to predict physical consequences, zero-shot generalize to new scenarios, and verify safety before execution — none of which LLMs alone can solve. This is the fundamental investment thesis for the track.

LLMs travel the road of "descriptions about the world"; World Models travel the road of "the world itself."

But for investors, this track has three special characteristics that must be confronted:

First, technical divergence greatly exceeds consensus. End-to-end vs. modular, the four World Model paths failing to converge, and whether Scaling Law applies — none of these have answers yet. Betting on a paradigm rather than a single company may be safer than picking winners.

Second, capital concentration is extreme. The top 6 companies capture the vast majority of leading capital, meaning this is a game where only a few players can get a seat. The valuation gap between mid- and late-stage companies and the leaders is an order of magnitude — there is no "catch-up" logic.

Third, deployment timelines may be far longer than LLMs. Sergey Levine says ten years; optimists say 18 months — this divergence itself is a signal. Short-term (2–3 years): invest in companies with strong industrial deployment capabilities, fundamentally not much different from the previous wave of robotics companies. Long-term (5+ years): focus on companies accumulating data flywheels in industrial scenarios with the ability to migrate to consumer scenarios.

Value is migrating from hardware to the model layer, but the moat for pure-model companies depends on data accumulation and depth of engineering integration. If the Generalist Policy paradigm works out, an ARM-like role will emerge; if end-to-end wins, integrated hardware-software companies benefit more; if neither materializes, big tech full-stack plays (especially Google and NVIDIA) will capture most of the value.

Appendix

Terminology & Market Size


A. Core Terminology

Core terminology
TermMeaning
VLM / VLAVision-Language Model / Vision-Language-Action Model
JEPAJoint Embedding Predictive Architecture — a joint embedding predictive architecture proposed by LeCun
Flow MatchingA generative method using optimal transport to map from noise to target distribution
Diffusion PolicyAn action-generation method based on diffusion models
UMIUniversal Manipulation Interface — a universal manipulation interface proposed by Columbia University
OXEOpen X-Embodiment — an open-source cross-embodiment robot dataset
Action ChunkingPredicting multiple future actions at once to reduce inference frequency requirements
Exhibit 10Source: Implic Capital research

B. Market Size (Goldman Sachs)

Humanoid robot market size and shipment forecasts
MetricData
2026 shipment forecast50,000–100,000 units
2030 shipment (base case)250,000+ units
2035 market size$38B
Unit cost trendDeclining to $15,000–20,000
2025 actual (Unitree + Agibot)~10,000 units
Exhibit 11Source: Goldman Sachs, Implic Capital research

Funding data: PitchBook, Crunchbase, public reporting, as of March 2026. This article is based on publicly available information, independently compiled by Implic Capital. It is for informational purposes only and does not constitute investment advice.