A new robot foundation model can learn physical tasks from a single demonstration lasting just 3 to 12 seconds, then attempt the task without any gradient updates or fine-tuning. Developed by Generalist AI, GEN-1.5 is designed to learn directly from physical examples rather than requiring engineers to retrain the model for every new job. The company says the model can infer what it is supposed to do from a short sensorimotor demonstration and immediately try the task itself. In tests across 10 physical tasks, GEN-1.5 achieved an average success rate of 59% when given just one demonstration. With five minutes of task-specific data and 10 gradient steps, its average success rate rose to 83%. The tasks were deliberately simple and short. They included twisting the lid off a glass jar, retrieving money from a purse, stacking cups, sweeping trash, opening a book, unzipping a pencil pouch and removing a vacuum pad. Robots learn by physical prompts GEN-1.5 is a large multimodal model that processes video along with sensor, language, and proprioceptive data. It maintains 30 seconds of information in context and generates action trajectories at 100 Hz. The demonstration itself acts as what Generalist AI calls a “physical prompt.” It can come from a person using handheld grippers or from the robot performing the task. Once the example is placed into the model’s context window, the robot attempts the task without a training phase. The company says it did not specifically train GEN-1.5 to perform in-context learning. It also did not add architectural changes, meta-learning loops or additional objectives designed to make the model improvise. That makes the reported result different from conventional robot training, where a new task can require substantial task-specific data and repeated optimization. GEN-1.5 can also combine physical prompts. In one demonstration, the model was given separate examples of unzipping a pencil pouch and retrieving money from it. It then connected the two behaviors into a continuous sequence, generating intermediate repositioning and recovery movements that were not present in either demonstration. One demo, different solutions The model also showed signs of generalizing beyond the exact actions it had seen. A demonstration recorded in simulation could be used to prompt a real robot, despite the model’s pretraining containing no simulated data. The resulting behavior could adapt to different hands and changes in object position and size. In another test, a person demonstrated a task with their own hands while visible to the robot’s cameras. The robot then reproduced the action using its own hands. Generalist AI also tested how the model responded when fine-tuned with only a few gradient steps. GEN-1.5 could adapt to new tasks in one to 10 steps using one to five minutes of data, equivalent to about 10 to 50 demonstrations. The model sometimes went beyond what it had been shown. After learning to use a brush to sweep a block into a bowl, it used a banana as an improvised brush when presented with one. Given a dustpan, it developed a different strategy, using the tool to lift and dump the block into the bowl. The company says these behaviors point toward a different approach to robot programming. Instead of writing detailed instructions for each new task, users could eventually demonstrate what they want and let the robot work out the physical details.Get the latest in engineering, tech, space & science - delivered daily to your inbox.With over a decade-long career in journalism, Neetika Walter has worked with The Economic Times, ANI, and Hindustan Times, covering politics, business, technology, and the clean energy sector. Passionate about contemporary culture, books, poetry, and storytelling, she brings depth and insight to her writing. When she isn’t chasing stories, she’s likely lost in a book or enjoying the company of her dogs.
New robot model learns physical skills from just 3–12 seconds of a single demo
Full Article
Original Source
Read the full article at Interestingengineering →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.