Generalist AI Develops Robots That Learn From Video
Robotics startup Generalist AI has built a model that allows robot arms to learn physical tasks on the spot from instructional videos, bringing highly adaptable automation closer to reality.

Massachusetts-based startup Generalist AI is developing a robotic model designed to understand the physical world and learn new tasks instantly. Founded by former Google DeepMind and Boston Dynamics researchers Pete Florence, Andrew Barry, and Andy Zeng, the company trains robot arms to perform chores like stacking cups and sorting blocks after watching a single instructional video. This zero-shot learning approach allows the machines to improvise, such as using a dustpan as a makeshift broom or switching grippers to unzip a purse when one angle fails.
Unlike traditional robotic training that requires feeding thousands of specific examples into a model—making them highly sensitive to minor environmental changes like lighting—Generalist AI builds its models from scratch. The startup gathers physical interaction data at a massive scale using custom, camera-equipped gloves that mimic robot pincers. Co-founder Pete Florence compared this capability to OpenAI's GPT-3, noting that a user could "prompt it to do a new task" and get immediate results.
For robotics practitioners, this shift from rigid, task-specific programming to generalized physical intelligence could revolutionize industrial deployment. However, the technology is still in its early stages. Generalist AI reports that its robots currently achieve a success rate of only about 59 percent, far below the 99 percent threshold required for reliable commercial deployment. Despite this limitation, experts note that the startup's hardware-agnostic data collection approach makes it one of the closest to achieving deployable, adaptable automation in real-world settings.
This is our own summary of reporting by WIRED AI



