METR · metr.org/time-horizons
pixels · motion · contact · a body moving through a room
"Agentic AI looks ready for the physical world."
Waddle Labs, July 2026
Real hardware, a competitor's footage, attributed.
Humanoid under model control · Anthropic
Pixels, forces, states. Text reasoning has nothing to perceive any of it with, and nothing to check itself against. That gap is physical AI.
Expensive. Not generalizable.
We can scrape the internet for videos, but how about actions?
[Google DeepMind] Genie 3
Trained on: blue
Tested on: red
[Google DeepMind] ALOHA Unleashed: A Simple Recipe for Robot Dexterity. Trained to fold a blue shirt, it generalizes only to a blue shirt.
DreamerV3 — one world model, many simulated domains, none of them the real one.
Ma et al., "Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild"
Generalizable. Works across all domains. Superhuman.
Predicts low-level latent appearance streams and high-level affordance trajectories in one model.
4D motion read out of unlabelled 2D video is the action.
A frame only teaches the model something when perception reads it right.
The agentic layer picks what to attempt, so data collects itself in the field.
Predicting pixels, latents, and 4D trajectories, all self-trained.
What would a World Model look like if we start from a real embodied agent acting in the real world? It has to have:
Actions are relative 3D pose deltas, independent of any specific hardware or control system.
Learns from failure and exploration equally, not only expert trajectories.
Given history frames and an action, anticipate the consequence. Change the action, the future changes.
Trained on real egocentric video and full-body motion capture, not synthetic physics engines. Sample actions → roll them out → keep the path closest to the goal across complex, uncontrolled environments.
Input












Goal
RGB · query points
lifted to 3D
Predicting pixels, latents, and 4D trajectories, all self-trained.
Reliable perception and action.
Five changes. You found them by looking back and forth — the encoder never gets to.
Every frame embedded alone.
Carries state across frames — what moved is legible in the embedding itself.
Sample output: "The small blue rubber sphere was moved." Read directly off the embedding, no extra reasoning step.
Did you realize the two images are different?
Five changes.
Two images.
And when you were comparing, did you eyeball the two images back and forth?
Current visual encoders in AI models can't.
Reference — Replace person located slightly off-center in the upper-middle section with a large rock.
Baseline — Remove the person standing on the rocky outcrop, arms raised, lower-middle of the image.
+ Ontos — Replace person positioned in the upper-central area of the image with a large rock.
Reference — Some detached houses appear beside the road on the bareland.
Baseline — The scene is the same as before.
+ Ontos — A building appears at the top of the scene.
Reference — The main image has additional findings of lung opacity and atelectasis vs. the reference image. It is missing the finding of cardiomegaly vs. the reference image.
Baseline — The main image has additional findings of atelectasis and cardiomegaly. It is missing the findings of pleural effusion and lung opacity.
+ Ontos — The main image has additional findings of atelectasis and lung opacity. It is missing the finding of cardiomegaly.
Understand the goal, plan long-horizon actions, use language and vision, predict the future.
React to contact, adjust force and trajectory, correct slip and errors, keep contact stable.
Ontos augments VLAs with a fast tactile reflex pathway. Slow deliberation sets the direction; fast reaction delivers the dexterity.
C³: Calibrated Conformal Control for Verifiably Steerable Foundation Models
Agentic harness for Physical AI.
[Google] Code-as-Policies: an LLM turns an instruction into a program built from perception and control primitives. Liang et al., ICRA 2023 · code-as-policies.github.io
A multimodal LLM–VLA that runs one module per modality, each at its own rate — vision, tactile, language, sound. Foveation, comparison, grounding.
A recursive world model built on a 4D-conditioned latent sequence model — and the 4D structure is self-trained from 2D pixel prediction, so no one has to label it.
An agentic robotic harness that imagines futures, forms plans, then executes, refines and banks the skill it just learned.
BDD100K · CVPR 20pixels, depth, motion
ARM4R · ICML 25vision, action, force
LISAt · NeurIPS 25SAR, multi-spectral, volumetric
CT Seg · arXiv 263D volumes, small changes
NWM · CVPR 25space and time, interactive
PEVA · NeurIPS 25continuous, everyday scenes
Two we already showed.
Sample efficient dexterous manipulation learning leading to self-exploration and learning. Carry stacked dishes and cups like human waiter. Use a Sharpa hand to manipulate chopsticks to pick up sushi, or an ice cube.
PEVA++ enabled cross embodiment mocap transfer: agile and on demand reconfigurable NPCs.
Reasoning over diverse, heterogeneous Geospatial data. We need to infer complex relationships across and between sensors to understand environmental risk.
Build easily, launch fast, and connect with millions in the new Ring Appstore.
Event-based notifications, workflow automation, customer traffic analysis.
Motion analysis, activity alerts, daily summaries.
Behavior pattern analysis, health insights, mood detection.
Pool condition monitoring, package theft detection.
Amazon Ring Marketplace API — home and office cameras. Pipeline, not signed.
AGPI — artificial general physical intelligence — is the merged model these three capabilities fuse into.




The first planet-scale capture system. Built by our CEO.
We figured out how to reason directly in raw sensory data: no language supervision, no curated labels. That unlocks a structural data advantage others cannot replicate.
Visual sentences · UVD-V1 · CVPR 2024
The team has shipped this class of model before — we get to a working system faster than anyone starting from scratch.
Not something an average team reproduces. The design choices come from having built the prior generation.
Open weights build the ecosystem and the funnel; the proprietary stack stays ours.
Fleets, labs and design partners already in reach — distribution and interaction data from day one.
Experience spread evenly across the four domains — two in each, so no single vertical carries the model.
Pipeline, not signed. Initials are who owns the relationship.
Two-year runway. Same plan — the delta is compute.
Primary ask $100M at $1B post. $180M is the same milestones with compute bought forward. · AMI · World Labs · Skild · Liquid AI · Physical Intelligence
FlexOlmo: Open Language Models for Flexible Data Use
There is no internet of physical interaction.
Plan, search, and backtrack still happen in text, not in the world.
The loop closes in milliseconds, not seconds.
Revenue lines are cumulative — the earlier ones keep running as the next one opens.
Fleet access, paid in model credits.
Sustained access, for a stake.
Offsets the $40M training line.
Paid deployment from day one.
NVIDIA · Panasonic · Bosch · Apple · Robot.com · EA · Kitware / NGA · Ring