Generalist AI's robots learn new tasks from a single video

Generalist AI's robots learn new tasks from a single video

Wired's Will Knight visited the Cambridge, Massachusetts offices of Generalist AI, a robotics startup about 15 minutes from his house, and watched robot arms perform simple chores such as stacking cups and putting blocks into bowls. He was struck by how quickly the arms figured out what to do: they mastered a range of tasks after ingesting a short instructional video, with no specific training given for that particular task.

Two demonstrations stood out. In one, a robot was told to sweep a block into a bowl using a dustpan and brush. When the brush was taken out of the scene, the robot improvised, using the dustpan alone like a brush to flick the block into the bowl. In another, a two-armed robot watched a video clip of someone unzipping a purse and removing banknotes, then did the same with a different purse. When it could not grab the money, it switched from its right gripper to its left to get a better angle. "Ha," said an engineer standing nearby. "It never did that before."

Generalist cofounder and CEO Pete Florence compared the moment to the early excitement around GPT-3, OpenAI's large language model released in 2020: "This is exactly the kind of thing people were really excited about with GPT-3. You could take that model and just prompt it to do a new task and it would have a real shot at doing it." Generalist is focused on teaching its robots the physics of the world, an approach the company believes helps a skill learned in one scenario transfer to another. Some of the demos evoked how children improvise when shown a task: researchers said they have often been surprised by what the robot decides to do, such as one robot that chose to sweep up items with a banana that had been placed in front of it.

Florence, Andrew Barry (cofounder and CTO), and Andy Zeng (cofounder and chief scientist) have all previously worked at Google DeepMind and Boston Dynamics on some of the most advanced hardware and robotic models around. The article contrasts their approach with traditional robot training, which feeds thousands of examples into a model, calling it a notoriously imperfect kind of learning: a robot trained that way will struggle if something as simple as the lighting changes.

Generalist builds special gloves resembling robot pincers, fitted with cameras, that people wear to perform different chores and generate training data; Knight saw a crate piled with several hundred of these gloves destined for workers in Mexico and elsewhere. Florence and team are cagey about the exact recipe used to train the robots, but say the company has already gathered a huge amount of high-quality training data. Unlike some other companies chasing smarter robots, Generalist has built its AI models entirely from scratch rather than relying on an open-source language model.

Outside roboticists who know the company's work spoke favorably of it. Danfei Xu of Georgia Tech said Generalist stands out among companies chasing more general robot models: "They have pushed this to the extreme, and they've done a really good job executing." Besides gathering a huge amount of high-quality data, he said, "they are excellent roboticists, and they have done really good science." Xu added that the demos suggest the company has an eye on deploying robots in real commercial settings, calling Generalist "the closest to something that's deployable." Karen Liu of Stanford University described Generalist's data approach as collecting physical interaction data at large scale without tying it too closely to one particular robot, and said "their strongest results suggest that this bet may be working."

Generalist itself says the learning skills of its models are not yet all that reliable: a robot completes a task it has already been shown correctly only about 59 percent of the time, on average, well short of the company's own goal of upwards of 99 percent. It also remains unclear how well the skills will generalize to every imaginable task or setting.

The piece closes with an anecdote from one recent evening: an engineer was stacking small cups on a table in front of a two-armed robot just to see what it would do. The robot joined in, grabbing and stacking other cups with its two grippers until it had built one neat pile, and the engineer began yelling in delight. This account is drawn from an edition of Wired's AI Lab newsletter, written by Will Knight.

Key facts

  • Generalist AI's robot arms learned to sweep a block into a bowl and unzip a purse to grab banknotes after watching a single instructional video, with no additional training for either task.
  • When conditions changed, the robots improvised: one used a dustpan alone to flick a block into a bowl after its brush was removed, and another switched from its right gripper to its left to get a better angle on banknotes it could not initially grab.
  • CEO Pete Florence compared the moment to GPT-3, saying you could prompt that model to try a new task and it would have a real shot at doing it.
  • Generalist collects training data through camera-equipped gloves worn by human workers; a Wired reporter saw several hundred of these crated for shipment to Mexico and elsewhere, and the company built its AI models entirely from scratch rather than on an open-source language model.
  • Generalist says its robots complete a task they have already been shown correctly only about 59 percent of the time on average, short of the company's own goal of upwards of 99 percent, and outside roboticist Danfei Xu still calls it "the closest to something that's deployable."

Why it matters

Traditional robot training feeds thousands of task-specific examples into a model, a process the article calls notoriously imperfect since a robot trained that way can struggle if something as simple as the lighting changes. Generalist's robots instead master a new chore after watching one short instructional video, with no separate training for that task. CEO Pete Florence frames this as a GPT-3-style moment: a model that can be prompted toward a new task and have a real shot at succeeding, rather than one that has to be trained for it. The improvisation on display, using a dustpan alone once the brush was gone, switching grippers to reach banknotes, sweeping with a banana placed in front of it, points to a model that has picked up something like physical intuition rather than memorized fixed motions.

Who it affects

Manufacturers and other businesses weighing flexible automation stand to gain if a robot can pick up a new task from a single demonstration instead of a lengthy, task-specific training process. It also affects the human workers Generalist hires to wear its camera-equipped gloves and record chores for training data, several hundred of which a Wired reporter saw crated for shipment to workers in Mexico and elsewhere. The approach sets Generalist apart from other companies chasing smarter robots that build on existing open-source language models rather than training from scratch, and it is being watched by outside roboticists such as Danfei Xu of Georgia Tech and Karen Liu of Stanford University, both of whom are already familiar with the company's work.

How to use it

There is nothing to use yet. The reporting describes a private demo at Generalist's own offices, not a released product: no product name, price, or availability is mentioned. Florence and team are described as cagey about the exact recipe behind the training process. Danfei Xu says the demos suggest Generalist has an eye on deploying robots in real commercial settings and calls it "the closest to something that's deployable," but that is an outside assessment of potential, not a statement that the system is available to buy, license, or deploy today.

How solid is it

The account comes from a Wired journalist who watched the demos in person, not from a controlled or published benchmark. Two roboticists outside the company, Danfei Xu at Georgia Tech and Karen Liu at Stanford, are described as familiar with Generalist's work and both speak favorably of its data approach and execution. Working against that, Generalist withholds the specifics of its training recipe, and the headline figures, the 59 percent task-completion rate and the description of a "huge amount" of training data, come from the company itself rather than an independent audit. The founders' backgrounds add some weight: Florence, Barry, and Zeng previously worked at Google DeepMind and Boston Dynamics on advanced hardware and robotic models, though the article does not say which founder worked at which lab.

Risks and caveats

By Generalist's own account, the system is not yet reliable: a robot completes a task it has already been shown correctly only about 59 percent of the time, on average, short of the company's stated goal of upwards of 99 percent. It also remains unclear how well the learned skills will generalize to every imaginable task or setting. The article gives no funding figures, valuation, headcount, or timeline for commercial deployment, and the evidence for the approach rests on a set of demos shown to a visiting journalist rather than systematic, independently verified testing.

“This is exactly the kind of thing people were really excited about with GPT-3. You could take that model and just prompt it to do a new task and it would have a real shot at doing it.”

— Pete Florence, Generalist AI cofounder and CEO