Skild AI has released S1, a robot foundation model designed to learn an unfamiliar physical task from one video demonstration. Instead of collecting a new dataset and retraining the model, an operator records the desired work and supplies the clip as a prompt. S1 then interprets the objects, intent and sequence before mapping them to the robot in front of it.
The company says the model can handle tasks lasting up to 10 minutes, including plant potting, coffee brewing and kit assembly. In one test, a team moved from recording a plant-potting demonstration to autonomous execution on hardware in 11 minutes. Skild reports about 66% success at each step on new multistep tasks, compared with 9% for a similar system, though these are company-run results rather than an independent evaluation.
S1 was developed using Nvidia infrastructure, including simulation and synthetic-data tools. Skild, Nvidia and Foxconn are also deploying the broader Skild Brain system on dual-arm robots assembling Nvidia Blackwell hardware.
The practical promise is faster adaptation on production lines where layouts and products change. The limitation is reliability: a 66% per-step success rate can compound across a long sequence, so real deployments still require validation, recovery systems and human oversight.