
Mechanic
Repairing an engine
- 01locate valve cover
- 02grasp rocker arm
- 03seat & align
- 04verify clearance
We turn real-world, first-person video of experts at work into structured, robot-ready training data — hands, tools, actions and 3D trajectories.

Prototype pipeline output on licensed stock footage · illustrative
Learning from experts across

Repairing an engine

Servicing industrial equipment

Preparing a meal

Assembling a product

Performing a procedure
Hover a card to see the demonstration it contains.
Our models transform raw egocentric video into structured representations designed specifically for training physical AI. Drag the divider.

"task": "loosen_bolt""domain": "automotive""subtasks": ["reach","grasp","seat","rotate"]"objects": ["ratchet_3/8","socket_12mm"]"hand": "right""grasp_type": "power_cylindrical""keypoints": [21, 3] // per frame"contact": { t: 15.2s, obj: ratchet }"tool_axis": [0.06, 0.99, 0.08]"rotation": -94.5° // ccw"trajectory": SE(3) × 212"success": trueHierarchical segmentation of long-horizon work into goals and steps.
Open-vocabulary detection and tracking of every tool and part in play.
21-point hand pose and full-body kinematics, per frame.
Contact events, grasp types and force-bearing moments.
Ordered action graphs with timing, pauses and retries.
Reusable skill primitives mapped to a shared taxonomy.
Metric 6-DoF paths of hands and tools lifted from monocular video.
Labeled outcomes, including the mistakes experts recover from.
We turn unstructured video into the building blocks of robot learning.
Head-mounted and body-worn footage from people doing real work.
Hands, bodies, objects and depth, reconstructed from every frame.
Long videos segmented into goals, subtasks and skill primitives.
Contacts, grasps and state changes grounded in 3D.
Robot-ready trajectories in standard learning formats.
Policies and world models pre-trained on human experience.
Instead of asking robots to collect every demonstration themselves, we learn from the humans who already know how to do the work.

Illustrative estimates from public reporting and our own early pilots. Not a guarantee of future performance.
The long-term vision is a foundational dataset of human interaction with the physical world — not captions, but complete, grounded procedures.
“A person is holding a screwdriver.”
“A person identifies the correct fastener, reaches for a screwdriver, grasps it, positions it against the fastener, applies torque, verifies the result, and moves to the next step.”
At scale, these trajectories become a library of reusable physical skills.
By learning from enormous collections of expert trajectories, we can work toward models that predict the consequences of physical actions.
fθ( state, action ) → future statelearned world modelRobots imagine the outcome of an action before executing it.
Learned environments let policies practice millions of times, safely.
Acquire skills that were never demonstrated directly on a robot.
Video becomes the foundation for a learned environment.
We're early. These are preliminary numbers from an internal pilot — small, honest, and improving every week.
score, 0–100 · higher is better · preliminary
Internal evaluation on 150 held-out egocentric clips across 6 task domains. Baseline: a general-purpose vision-language model prompted zero-shot. Results are preliminary and not peer-reviewed.
A technical founding team working at the intersection of computer vision, human motion and robot learning.
Robotics & computer vision. Previously building perception systems for autonomous vehicles.
Robotics teams and physical-AI labs that need large, diverse demonstration data to pre-train manipulation policies and world models — without standing up a fleet of teleoperated robots.

We're partnering with robotics teams, data contributors and investors who want to build the foundation model for physical work.
alexyang@electric-ai.link