ElectricAI LabsThe data layer for physical AI

The world's experts are already generating the data robots need.

We turn real-world, first-person video of experts at work into structured, robot-ready training data — hands, tools, actions and 3D trajectories.

hours of egocentric video processed in pilot
1,200+
skill primitives in our taxonomy
40+
annotation layers per frame
9
electric-ai / pipeline / session_0412 / cam_ego_01.mp4
First-person view of a mechanic's hands tightening a bolt with a wrench
wrench_19mm · 0.97
hand_R · power_grasp
REC00:16.8
reach
grasp
align
torque
verify

Prototype pipeline output on licensed stock footage · illustrative

Learning from experts across

  • Automotive repair
  • Industrial maintenance
  • Food preparation
  • Electronics assembly
  • Surgical training
  • Warehouse logistics
  • Machining
  • Construction trades
  • Lab automation
  • Home services
  • Automotive repair
  • Industrial maintenance
  • Food preparation
  • Electronics assembly
  • Surgical training
  • Warehouse logistics
  • Machining
  • Construction trades
  • Lab automation
  • Home services
01 · The insight

Every expert is a robot demonstration.

A mechanic repairing an engine. A chef preparing a meal. A surgeon performing a procedure. These people aren't just doing their jobs — they are generating demonstrations of how to interact with the physical world.
Mechanic: Repairing an engine
hand_R · grasp
01 / 05

Mechanic

Repairing an engine

  1. 01locate valve cover
  2. 02grasp rocker arm
  3. 03seat & align
  4. 04verify clearance
Technician: Servicing industrial equipment
caliper · 0.94
02 / 05

Technician

Servicing industrial equipment

  1. 01pick up caliper
  2. 02open jaws
  3. 03measure bore
  4. 04log reading
Chef: Preparing a meal
knife · 0.98
03 / 05

Chef

Preparing a meal

  1. 01stabilize produce
  2. 02position blade
  3. 03slice ×12
  4. 04transfer to tray
Assembler: Assembling a product
iron_tip · 0.92
04 / 05

Assembler

Assembling a product

  1. 01pick component
  2. 02orient to pad
  3. 03insert
  4. 04solder & inspect
Surgeon: Performing a procedure
forceps · 0.95
05 / 05

Surgeon

Performing a procedure

  1. 01request instrument
  2. 02handoff
  3. 03precision grip
  4. 04return

Hover a card to see the demonstration it contains.

02 · What we extract

From raw pixels to robot-ready structure.

Our models transform raw egocentric video into structured representations designed specifically for training physical AI. Drag the divider.

Annotated video frame: hand keypoints, tool detections and rotation axis
hand_R · power_grasp
ratchet_3/8 · 0.96
socket_12mm · 0.93
engine_bay · ctx
STRUCTURED
RAW VIDEO
demo_000184.json● valid
"task": "loosen_bolt""domain": "automotive""subtasks": ["reach","grasp","seat","rotate"]"objects": ["ratchet_3/8","socket_12mm"]"hand": "right""grasp_type": "power_cylindrical""keypoints": [21, 3]  // per frame"contact": { t: 15.2s, obj: ratchet }"tool_axis": [0.06, 0.99, 0.08]"rotation": -94.5°  // ccw"trajectory": SE(3) × 212"success": true
01

Tasks & subtasks

Hierarchical segmentation of long-horizon work into goals and steps.

02

Objects & tools

Open-vocabulary detection and tracking of every tool and part in play.

03

Hand & body motion

21-point hand pose and full-body kinematics, per frame.

04

Object interactions

Contact events, grasp types and force-bearing moments.

05

Temporal sequences

Ordered action graphs with timing, pauses and retries.

06

Actions & skills

Reusable skill primitives mapped to a shared taxonomy.

07

3D trajectories

Metric 6-DoF paths of hands and tools lifted from monocular video.

08

Success & failure

Labeled outcomes, including the mistakes experts recover from.

03 · The pipeline

From video to robot intelligence.

We turn unstructured video into the building blocks of robot learning.

  1. 01

    Expert Video

    Head-mounted and body-worn footage from people doing real work.

  2. 02

    Perception

    Hands, bodies, objects and depth, reconstructed from every frame.

  3. 03

    Tasks & Skills

    Long videos segmented into goals, subtasks and skill primitives.

  4. 04

    Actions & Interactions

    Contacts, grasps and state changes grounded in 3D.

  5. 05

    Structured Demonstrations

    Robot-ready trajectories in standard learning formats.

  6. 06

    Robot Training

    Policies and world models pre-trained on human experience.

04 · The data layer for physical AI

Robots shouldn't have to generate all their own training data.

Instead of asking robots to collect every demonstration themselves, we learn from the humans who already know how to do the work.

Today

Robot teleoperation

  • Limited in scale
  • Expensive to collect
  • Tied to specific hardware & environments
What we're building

Human expert video

  • Massive in scale
  • Naturally diverse
  • Captured across environments, tools & tasks

Order-of-magnitude comparison

Human expert video Robot teleoperation
Cost per hour of demonstration
lower is better
< $5
$100–$200
Distinct environments per 1k hrs
higher is better
1,000+
~10
Hours collectable per month
higher is better
100k+
~1k

Illustrative estimates from public reporting and our own early pilots. Not a guarantee of future performance.

05 · The vision

Millions of hours. Thousands of skills.

The long-term vision is a foundational dataset of human interaction with the physical world — not captions, but complete, grounded procedures.

Not just

“A person is holding a screwdriver.”

But

“A person identifies the correct fastener, reaches for a screwdriver, grasps it, positions it against the fastener, applies torque, verifies the result, and moves to the next step.”

action_graph · fasten_panelstep 0/7
gaze
fixate
track
inspect
hand_R
reach
grasp
position · torque
release
hand_L
stabilize part
tool
screwdriver_PH2
t0t1t2t3t4t5t6

At scale, these trajectories become a library of reusable physical skills.

Pick up.
pick_up()align()insert()rotate()connect()tighten()inspect()assemble()repair()
06 · Beyond imitation

Not copying humans. Understanding how actions change the world.

By learning from enormous collections of expert trajectories, we can work toward models that predict the consequences of physical actions.

peg_A · grasped
statesₜ
+
Δpose · lift → move → insertΔz −24mm
actionaₜ
peg_A · insertedp = 0.93
future stateŝₜ₊₁
fθ( state, action ) future statelearned world model

Reason before acting

Robots imagine the outcome of an action before executing it.

Train in simulation

Learned environments let policies practice millions of times, safely.

Learn without demos

Acquire skills that were never demonstrated directly on a robot.

Video becomes the foundation for a learned environment.

07 · Early results

Early signal from our prototype pipeline.

We're early. These are preliminary numbers from an internal pilot — small, honest, and improving every week.

0+
Hours of egocentric video processed
0+
Skill primitives in our taxonomy
0%
Subtask segmentation F1
0 mm
Mean 3D hand-pose error

Internal benchmark · 150 held-out expert clips

score, 0–100 · higher is better · preliminary

ElectricAI pipeline (v0.3) Off-the-shelf VLM baseline
Subtask segmentation F1
91
64
Tool identification accuracy
94
78
Contact-event recall
87
52
Grasp-type classification
82
49

Internal evaluation on 150 held-out egocentric clips across 6 task domains. Baseline: a general-purpose vision-language model prompted zero-shot. Results are preliminary and not peer-reviewed.

08 · Team

A small team, obsessed with physical AI.

A technical founding team working at the intersection of computer vision, human motion and robot learning.

Chinmay Govind

Chinmay Govind

Co-founder & CEO

Robotics & computer vision. Previously building perception systems for autonomous vehicles.

Alex Yang

Alex Yang

Co-founder & CTO

Machine learning for video understanding and human motion capture.

We're hiringResearch engineers in vision, 3D & robot learning
09 · FAQ

Questions we hear often.

Robotics teams and physical-AI labs that need large, diverse demonstration data to pre-train manipulation policies and world models — without standing up a fleet of teleoperated robots.

Human expertise → Data → Intelligence → Robots

The physical world is the next training corpus.

We're partnering with robotics teams, data contributors and investors who want to build the foundation model for physical work.

alexyang@electric-ai.link