Skild AI is teaching robots new tasks from a single video — and NVIDIA showcased the capability on its blog as part of a robotics-heavy September that also featured the world’s robotaxi leaders building on NVIDIA’s stack.
From Teleoperation to One-Shot
Training a robot a new task traditionally means teleoperating it through hundreds of demonstrations, then training a task-specific policy. Skild’s approach — running on NVIDIA’s physical AI infrastructure — collapses that: show the robot a video of a task once, and the S1 foundation model generalizes it to the robot’s body and environment.
Single-video (one-shot) imitation is the robotics analogue of what large language models did for text skills: the bet is that a general foundation model, trained on broad data, can pick up specific tasks from minimal examples. Generalization in physical AI remains the field’s open problem — one that Luma’s newly announced Open Physical AI Lab is also attacking.
The Robotics Moment
The demonstration joins a wave: NVIDIA’s Vera CPUs landing at AI labs, Figure’s human-vs-robot showcase, Enigma’s interactive robots, and physical AI taking the wheel at robotaxi companies. Read the full story at blogs.nvidia.com/blog/skild-ai-s1-physical-ai.