Generalist AI's Robot Learns New Tasks from Video Clips Alone
Original: I Saw the Future of AI in a Robot That Can Learn on the Spot
Why This Matters
Generalizable physical AI could unlock scalable robotics deployment across unstructured real-world environments.
Cambridge, MA-based startup Generalist AI demonstrated robotic arms that master new tasks after watching a short instructional video—with no task-specific training. Robots improvised mid-task, including using a banana as a sweeping tool, surprising even the engineers on site.
Generalist AI, a Cambridge, Massachusetts-based robotics startup, showcased robotic arms capable of learning and adapting to new tasks in real time by watching brief instructional video clips—requiring no task-specific training data. WIRED reporter Will Knight visited the company's offices and observed several striking demonstrations. In one, a robot instructed to sweep a block into a bowl improvised by using a dustpan as a brush when the actual brush was removed. In another, a two-armed robot watched a video of someone unzipping a purse and removing banknotes, then replicated the action with a different purse—switching grippers mid-task to find a better angle, a behavior its engineers had never seen before. In yet another unscripted moment, a robot chose to use a banana as a sweeping tool when it was placed in its environment. CEO Pete Florence compared the model's generalization capability to GPT-3's ability to perform new tasks from prompts alone. The company focuses on teaching robots physical intuition—an approach inspired by how human infants learn about the world. Cofounders Florence, Andrew Barry (CTO), and Andy Zeng (Chief Scientist) all previously worked at Google DeepMind and Boston Dynamics. Generalist AI uses custom gloves with cameras to collect human demonstration data, aiming to build a general-purpose robotic foundation model rather than narrowly trained task-specific systems.