Track & Funding: The Inflection Point Has Arrived, Concentration Is Extreme
1.1 What Is a World Model?
Any technical direction that attempts to have AI systems build an internal representation of the physical world, predict the consequences of actions, and thereby achieve generalized robot control can be classified under the World Model umbrella. It has two interpretations — narrow and broad:
| Category | Representative Directions | Core Idea |
|---|---|---|
| Narrow World Model | JEPA (LeCun), video prediction models | Predict future states in latent or p… |
| Spatial Intelligence | World Labs (Fei-Fei Li) | Reconstruct 3D world from 2D ima… relationships |
| Broad Embodied AI | VLA, Generalist Policy, Hierarchical Systems | All foundation-model approaches … perception to action |
1.2 Funding Landscape: More Crowded Than LLM in 2015
Inflection points (2023H2 → 2024H1):
- RT-2 paper (July 2023): Demonstrated that VLMs can directly control robots
- OpenAI returns to robotics (2023–2024): Invested in 1X and Figure, lending credibility to the sector
- Figure $675M mega-round (February 2024): Bezos + Microsoft + NVIDIA + OpenAI at the same table, setting the valuation anchor
- NVIDIA GTC unveils GR00T (March 2024): Jensen Huang declares humanoid robots the next platform
- Pi $400M round (November 2024): Validates the standalone "Robot Foundation Model" track
The driving force is capital overflow from the LLM track — robotics is one of the largest capital-capacity tracks after large language models. But when exactly the inflection point will arrive is a matter of extreme disagreement within the industry: Pi CEO Sergey Levine says at least ten years before seeing large-scale deployment; Sunday CEO Tony Zhao says 18 months; some World Model company CEOs are as optimistic as 8 months.
1.3 Key Players and Segmentation
Companies in the space can be understood along two dimensions: which use case they enter from (robot manipulation vs. autonomous driving vs. general 3D), and which layer of the technology stack they operate in (world understanding layer vs. decision execution layer vs. hardware layer).
A notable trend: value is migrating from the hardware layer to the model layer. Figure has the highest valuation ($39B), but its core narrative has already shifted from "making robots" to "robot operating system"; Skild and Generalist make no hardware whatsoever and purely sell models/software. This mirrors the PC era, where hardware companies (Dell) gave way to operating systems (Microsoft).
Value is migrating from the hardware layer to the model layer — a replay of the PC era, where Dell gave way to Microsoft.
A. General-Purpose World Models
| Company | Valuation | Funding | Approach |
|---|---|---|---|
| AMI Labs | $3.5B | $1.03B Seed | JEPA |
| World Labs | $5B+ | $1.23B (incl. Feb 2026 $1B raise) | Spatial Intelligence / 3D |
| Odyssey | Undisclosed | $27M | General World Model… |
| Decart | $3.1B | $153M | Real-time generative … world model |
B. Robot Foundation Models
| Company | Valuation | Funding | Approach |
|---|---|---|---|
| Skild AI | $14B | $1.83B | Generalist Policy |
| Physical Intelligence (Pi) | $5.6B → in talks at $11B+ | ~$1.1B+ | VLA + Flow Matching |
| Sunday Robotics | $1.15B | ~$200M | Diffusion Policy + crowdsourcing |
| Generalist | ~$3B | $140M | Universal robot control |
C. Big Tech Landscape
Big tech is the most important variable in the space — theoretically best positioned, but with very different objectives:
- NVIDIA builds platforms, not robots. Cosmos and GR00T N1 are primarily GPU showcases, but open-sourcing GR00T N1 compressed the differentiation space for pure-model startups.
- Google DeepMind is building a full-stack in-house (Gemini Robotics, Genie 3, RT-2/RT-X). Demis Hassabis strongly believes World Models are the core path to AGI, posing the greatest threat to startups. Primary constraint: internal resource allocation.
- Amazon is scenario-driven and has acquired the Covariant AI team (~$200M); it is the largest robotics customer.
- OpenAI enables rather than builds — invested in 1X and Figure; its own robotics team remains in frontier research mode.
- Tesla builds in a closed ecosystem. Optimus Gen 3 has 1,000+ internal deployments, but progress is far behind public announcements.
D. Integrated Hardware-Software / Humanoid Robots
The source sheet attached to this segment duplicates the end-to-end vs. modular architecture comparison rebuilt in Section II (Exhibit 5); representative integrated hardware-software players named across this report include Figure, 1X, Tesla and Agility Robotics.
Technology Roadmap: More Divergence Than Consensus
Conclusion: Technical roadmaps have not converged, and will not converge in the near term. Three core points of divergence:
- End-to-end vs. modular: End-to-end has been achieved for single-scenario/simple tasks; complex scenarios remain a breakthrough challenge.
- Does Scaling Law apply to robotics: VLA scaling returns are not as clean a curve as LLMs — doubling model parameters doesn't double success rates. VLAs are constrained by a ~8B parameter ceiling (robot control requires ~35ms response latency).
- Is World Model a necessary condition: As of March 2026, no World Model system has beaten VLA baselines on complex manipulation tasks. JEPA papers look impressive; the engineering hasn't materialized yet.
2.1 Training Framework: Pre / Mid / Post-train
Robot model training is divided into three stages, each with different data requirements:
- Pre-train: Use internet video, game, and simulation data to build basic physical cognition. Equivalent to the physical intuition humans accumulate from "watching the world" since childhood.
- Mid-train: Use cross-embodiment robot datasets (Open X-Embodiment) and human motion data (UMI) to build generalized action understanding. Like having watched many trade videos but never done the work yourself.
- Post-train: Teleoperation data ($50–200/hour) to make the model work reliably on specific hardware and specific tasks. Simulation data is largely ineffective at this stage.
2.2 System Architecture: End-to-End vs. Modular
| Architecture | Technical Logic | Representative Companies |
|---|---|---|
| End-to-End | A single model from pixels to actions; lets the model learn its own intermediate representations, avoiding human-designed information bottlenecks | Pi (pi0, 3B params), Google Gemini Robotics, Sunday |
| Modular (Hierarchical) | Large upper-layer model for planning (0.1–2 Hz); small lower-layer model for execution (100–1000 Hz) | Figure, Skild, 1X, Agility |
Generalist Policy is the most aggressive sub-direction of the modular approach: training a single, cross-task, cross-embodiment general execution model (Skild is the representative). If it succeeds, it will play a role similar to ARM in the chip space. The risk: if the upper layer uses a general LLM and the lower layer uses scenario-specific small models, anyone can build this stack — the moat depends on data accumulation and depth of engineering integration.
2.3 Action Generation Methods: From Imitation to Generation
Whether end-to-end or modular, the model ultimately needs to output joint actions. This "final step" has gone through three generations of evolution in three years:
- Behavior Cloning (BC): Directly imitates human demonstrations. The problem is that the same task usually has multiple valid approaches (folding a shirt can go left-to-right or right-to-left), and BC can only learn the "average" approach — doing neither well. RT-2 (2023) used this method.
- Diffusion Policy: Borrowed from AI image-generation diffusion models, starting from noise and generating action trajectories through 10–100 denoising steps. Can express multiple valid approaches simultaneously, solving the BC "averaging" problem. Action chunking predicts an entire future action sequence at once (e.g., 16 joint angles over 1 second), reducing inference frequency requirements. Sunday co-founder Cheng Chi is a core author.
- Flow Matching: Uses optimal transport theory to find a "shortest path" direct mapping from noise to action, outputting in 5–10 steps — 5–10× faster than Diffusion. Pi's pi0 (2024) uses this method; currently considered the relatively leading approach.
2.4 Four Paths for World Models
| Path | Core Idea | Representative |
|---|---|---|
| JEPA | Instead of predicting pixels, predict abstract representations of future states in latent space. Focuses only on causal structure, ignores irrelevant details | AMI Labs (LeCun + Xie Saining, $3.5B) |
| Video Generation | Directly generate future video frames — essentially using video generation as a physics simulator | NVIDIA Cosmos (900 trillion token training), Google Genie 3, Decart |
| 3D Spatial | Instead of predicting the future, reconstruct a persistent 3D scene … | n/a |
JEPA / Video Gen focuses on the temporal dimension ("what happens next"); 3D Spatial Intelligence focuses on the spatial dimension ("what does the world look like"). The two are naturally complementary: first use 3D reconstruction to understand scene structure, then use temporal prediction models to plan actions.
2.5 World Model Companies and Embodiment
| Mode | Representative | Logic |
|---|---|---|
| Tightly Coupled (Hardware-Software Integrated) | Figure, 1X, Tesla | End-to-end requires deep software-hardware coupling; own hardware = own data |
| Pure Software (Hardware-Agnostic) | Pi, Skild | Build universal models/data to serve all hardware vendors; asset-light, high leverage |
This is fundamentally determined by technical approach: companies pursuing end-to-end VLA are theoretically able to be pure software; those pursuing Hierarchical with lower-layer Policy tightly coupled to hardware are not. Notably, pure-software companies are also beginning to build their own hardware — primarily as data collection vehicles (e.g., Sunday's home robot collects data in real user environments through product sales) — but their core DNA remains non-hardware.
Data: A "Good Vantage Point" Before Technical Convergence
When paradigms have yet to converge, data is the most valuable lens for observing this space — because regardless of which technical path wins, data requirements will continue to grow as new use cases unlock. This is a structural demand, not a one-time event.
3.1 Three-Stage Data Requirements
| Stage | Core Data | Volume & Cost |
|---|---|---|
| Pre-train | Internet video, games, simulation | Large volume; near-zero cost |
| Mid-train | Open X-Embodiment (21 institutions, 22 robot types, 527 skills), UMI human motion | Moderate |
| Post-train | Teleoperation data | $50–200/hour; tightly coupled to hardware |
Core tension: Pre-train data volume is sufficient, but quality (especially action alignment) is not. Sergey Levine estimates it will take 10 years to see a "GPT moment" level of emergence.
Simulation data effectiveness decreases across stages: high for Pre-train, moderate for Mid-train, essentially zero for Post-train — unless simulation physics accuracy sees an order-of-magnitude improvement, which will not happen in the near term. This gap has also spawned a unicorn: Lightwheel, tightly integrated with NVIDIA.
3.2 The Dilemma of Data Vendor Business Models
Data companies face a structural contradiction: low standardization (every client has different hardware and data formats) but limited client budgets and long decision chains. This is not a business that can scale quickly like SaaS — it is more like high-end consulting.
Representative vendors by category (international + Chinese):
Teleoperation
- Lightwheel AI — deeply integrated with NVIDIA; MimicGen generates 100–1000× synthetic data from a single demonstration
- XDOF.ai — tiered data system; clients include Pi, Amazon, Figure
- Sensei — crowdsourced teleop; hardware 10× cheaper than industry average
First-Person Perspective
- Build AI — open-source Egocentric-100K
- Asimov AI — self-operated cleaning company as data engine
- GenRobot — daily capture at ten-thousand-hour scale; 70%+ overseas revenue
Tactile / Motion Capture
- Pacini — world's top shipping volume for tactile sensors
- Noitom — 70%+ global share in inertial motion capture
- LUMOS — UMI leader
- Spirit AI — low-cost exoskeleton tactile glove, 21 degrees of freedom
Data Platforms / Simulation
- Galaxea — open-source G0 with 400K+ downloads; serves World Labs, Pi
- Zhiyuan — 4,000㎡ data factory; 30–50K trajectory records/day average
- Lightwheel — simulation tightly integrated with NVIDIA
- Palatial — physically accurate 3D simulation
Track assessment: The real opportunity lies in end-to-end solutions (collection → labeling → training → iteration, full pipeline). Raw data volume alone is not a moat; data quality, scenario coverage, and model adaptation capability are the true differentiators.
Commercialization: B2B Barely Works, B2C at Least 5 Years Away
4.1 B2B: Currently Only Warehouse Logistics Is Scaling
Core assessment: Warehouse logistics is the scenario that is currently showing initial traction, but the essence is "replacing existing automation with a more expensive solution" — ROI has yet to be proven.
The most representative case is Agility Robotics + Amazon: Digit was tested at Amazon's Sumner warehouse for 18 months, achieving a 98% task success rate and operating costs of $10–12/hour (vs. $30/hour for human labor). It has been in full-time deployment at GXO Flowery Branch for a year, cumulatively moving over 100,000 totes. Agility completed a $400M Series C in 2025, with Salem factory production capacity targeting 10,000 units/year.
Automotive manufacturing is ramping up: BMW + Figure (Spartanburg factory), Mercedes-Benz + Apptronik (Apollo factory testing).
Key bottlenecks: Technically, generalization is insufficient ("passing a demo doesn't mean it works on the production line"); on the sales side, decision chains are long (robots are capex decisions; large-client sales cycles are 6–12 months); budget-wise, potential clients have limited budgets, while large clients have enough but impose strict requirements.
4.2 B2C: Won't Truly Open Before 2030
All current "home robot" narratives are fundraising stories. Key unlock conditions:
| Condition | Current Status | Expected Unlock | Expected Timeline |
|---|---|---|---|
| Safety certification framework | Almost nonexistent | 2028–2030 | 2024–2026 |
| Cost < $10K | Currently $50K+ (high-end) / $16K (low-end) | 2030+ | 2025–2028 |
| Operational reliability > 99.9% | ~80–90% (lab setting) | 2030+ | 2027–2030 |
| Scenario generalization | Limited to controlled environments | 2028–2030 (limited scenarios) | n/a |
The most likely path: industrial scenarios (2024–2028) → commercial service scenarios (hotels, hospitals, 2028–2030) → home scenarios (2030+).
Investment Perspective: A Game for the Few
Returning to the "north slope vs. south slope" metaphor at the opening — LLMs travel the road of "descriptions about the world"; World Models travel the road of "the world itself." Robots need to predict physical consequences, zero-shot generalize to new scenarios, and verify safety before execution — none of which LLMs alone can solve. This is the fundamental investment thesis for the track.
LLMs travel the road of "descriptions about the world"; World Models travel the road of "the world itself."
But for investors, this track has three special characteristics that must be confronted:
First, technical divergence greatly exceeds consensus. End-to-end vs. modular, the four World Model paths failing to converge, and whether Scaling Law applies — none of these have answers yet. Betting on a paradigm rather than a single company may be safer than picking winners.
Second, capital concentration is extreme. The top 6 companies capture the vast majority of leading capital, meaning this is a game where only a few players can get a seat. The valuation gap between mid- and late-stage companies and the leaders is an order of magnitude — there is no "catch-up" logic.
Third, deployment timelines may be far longer than LLMs. Sergey Levine says ten years; optimists say 18 months — this divergence itself is a signal. Short-term (2–3 years): invest in companies with strong industrial deployment capabilities, fundamentally not much different from the previous wave of robotics companies. Long-term (5+ years): focus on companies accumulating data flywheels in industrial scenarios with the ability to migrate to consumer scenarios.
Value is migrating from hardware to the model layer, but the moat for pure-model companies depends on data accumulation and depth of engineering integration. If the Generalist Policy paradigm works out, an ARM-like role will emerge; if end-to-end wins, integrated hardware-software companies benefit more; if neither materializes, big tech full-stack plays (especially Google and NVIDIA) will capture most of the value.
Terminology & Market Size
A. Core Terminology
| Term | Meaning |
|---|---|
| VLM / VLA | Vision-Language Model / Vision-Language-Action Model |
| JEPA | Joint Embedding Predictive Architecture — a joint embedding predictive architecture proposed by LeCun |
| Flow Matching | A generative method using optimal transport to map from noise to target distribution |
| Diffusion Policy | An action-generation method based on diffusion models |
| UMI | Universal Manipulation Interface — a universal manipulation interface proposed by Columbia University |
| OXE | Open X-Embodiment — an open-source cross-embodiment robot dataset |
| Action Chunking | Predicting multiple future actions at once to reduce inference frequency requirements |
B. Market Size (Goldman Sachs)
| Metric | Data |
|---|---|
| 2026 shipment forecast | 50,000–100,000 units |
| 2030 shipment (base case) | 250,000+ units |
| 2035 market size | $38B |
| Unit cost trend | Declining to $15,000–20,000 |
| 2025 actual (Unitree + Agibot) | ~10,000 units |
Funding data: PitchBook, Crunchbase, public reporting, as of March 2026. This article is based on publicly available information, independently compiled by Implic Capital. It is for informational purposes only and does not constitute investment advice.