Dyna Robotics trains its robots on 1 million hours of human video

RoboticsAI Agents
Illustration generated by AI: Editorial image for Dyna Robotics trains its robots on 1 million hours of human video

The Core · TL;DR

  • Dyna Robotics' new Dyna-2 model was pre-trained on over 1 million hours of egocentric human video instead of teleoperation data
  • Quality pass rate jumped to 87% from 46% for predecessor Dyna-1, which already runs in hotels, restaurants, and laundromats
  • Scaling pre-training data alone pushed manufacturing task success from ~20% to as high as 90%, with post-training data unchanged
  • Just 13 minutes of robot-specific data taught five-fingered hands to open a bottle cap; the model is vendor-operated only, with no public weights or API

Dyna Robotics has released Dyna-2, a robot control model built almost entirely from footage of humans doing everyday tasks rather than from robot teleoperation data. The California-based startup, co-founded by Jason Ma, says the model was pre-trained on more than 1 million hours of egocentric human video before any robot-specific fine-tuning began.

That shift in training strategy shows up directly in performance numbers. In one customer deployment, Dyna-2 hit an 87% quality pass rate, nearly double the 46% recorded by its predecessor, Dyna-1, which already runs in production settings including hotels, restaurants, and laundromats.

Scaling pre-training, not just data volume

The team trained on nested subsets of exactly 1,000, 10,000, 100,000, and 1,000,000 hours of video to isolate the effect of scale. On manufacturing tasks, success rates climbed from roughly 20% to as high as 90% as the pre-training set grew, with the post-training data held constant throughout.

The same pattern held across 15 benchmark tasks, where accuracy consistently rose alongside the amount of human-video pre-training. Video co-training also boosted instruction-following scores by 133% on tasks that required the robot to change its movements based on different user commands.

Perhaps the most striking result involves data efficiency after pre-training. Researchers say just 13 minutes of robot-specific footage was sufficient to teach a pair of five-fingered robotic hands to twist open a bottle cap, a task that would typically demand far more targeted demonstration data.

Architecture and post-training tasks

Dyna-2 runs on a video-diffusion backbone paired with a mixture-of-transformers architecture, and it was validated across stationary robot arms, humanoid prototypes, and five-fingered hands. Post-training focused on practical, repeatable jobs: clearing trash trays, assembling first-aid kits, building totes, scooping food, tying rope, prepping hangers, and retrieving specific drinks from a fridge.

Unlike many recent robotics models released with open weights, Dyna-2 is only available as a vendor-operated system. There is no public checkpoint, API, or downloadable weight release, which keeps the model tightly coupled to Dyna Robotics' own deployment pipeline.

Manufacturing task success rates rose from about 20% to 80-90% purely by scaling human-video pre-training, without touching the post-training dataset.

The result reframes what counts as useful robot training data. If a model can absorb general physical competence from ordinary human video and then specialize with minutes rather than hours of robot demonstrations, the bottleneck for deploying robots in new environments looks a lot smaller than it used to.

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram