Site icon aivancity blog

Skild AI's S1: How Video Learning Is Truly Changing Robotics

Unveiledby Skild AI in August 2026, S1 is a foundational model for robotics designed to perform a new task based on a single video demonstration, without additional task-specific training. Published tests show progress on manipulation sequences lasting up to ten minutes, but they are still based on internal evaluations. The challenge therefore goes beyond the technical demonstration: it is to determine whether this context-based learning can become reliable, safe, and controllable enough for professional environments.

01

What Just Happened

Skild AI unveiled S1 in August 2026 as its flagship foundation model for robotic manipulation. The system is fed a video showing a task and then attempts to replicate the observed task using a robot, without modifying its weights or requiring a specific post-training phase. On September 10, NVIDIA detailed the infrastructure used to train, simulate, and deploy this model, including Cosmos, Omniverse, Isaac Sim, and Isaac Lab.[1][2]

The promise involves new and time-consuming tasks. Skild AI shows S1 repotting a plant, making filter coffee, making pancakes, or assembling a kit. Some sequences last up to ten minutes and involve several successive gestures. For repotting, the company reports that it went from recording the demonstration to beginning autonomous execution in eleven minutes.[1][2]

This timeline needs to be clarified: S1 was not launched on September 10. The initial announcement was made in August 2026, while NVIDIA’s September 10 release highlights the technical collaboration and the applications of its computing and simulation platforms.

02

What's Really New

The idea of teaching a robot a task through demonstration is not new. Demonstration-based programming, imitation, and teleoperation have been studied for several decades. What sets S1 apart is its ambition to transfer to robotics a mechanism that has become central to major language models: context-based learning.

In this context, the video is not used to retrain the model. It is placed in context at the time of use. S1 must extract the intention from it, recognize relevant objects, track the task’s progress, and translate the demonstration into actions compatible with the robot and the current scene. The company claims that the same set of weights generates all the examples shown.[1]

The claimed innovation also relates to the duration and novelty of the tasks. Skild AI notes that several scenarios did not appear in the pre-training data and that they combined basic gestures into novel sequences. If this capability is confirmed outside of internal testing, it could reduce the cost of data collection and specialization for each new operation.

MSc in Generative,
, and Agent-Based AI

Become an expert in generative and agent-based AI: LLMs, transformers, and enterprise deployment. Includes a learning trip to Silicon Valley. RNCP Level 7 certification (equivalent to a 6-year post-secondary degree).

12 months — 6 years of post-secondary education Master's degree (Bac+5) with experience Part-time — Fridays & Saturdays Paris-Villejuif Campus
Learn more about the program → RNCP Level 7 Certification
03

How S1 Learns from a Video

S1 is trained using episodes in which a visual demonstration describes the task to be performed. The model learns to relate what it observes in the video to the current scene and the available motor controls. It therefore does not necessarily copy every movement of the demonstrator. Instead, it seeks to reconstruct an action policy that achieves the same goal using its own robotic body.

Pre-training combines several types of data: teleoperation, first-person human videos, simulation, and experiments conducted on robots. Each addresses a different constraint. Teleoperation is similar to real hardware but expensive to produce. Human video provides greater diversity, with a wider gap between human gestures and those of the robot. Simulation is easier to develop, but it must accurately represent contact, pressure, collisions, and physical properties.[1]

NVIDIA is involved in several stages of this process. Cosmos helps organize and enrich video data; Omniverse and Isaac Sim provide virtual environments; Isaac Lab is used for reinforcement learning; and TensorRT optimizes inference. These tools do not constitute S1 on their own. They form the infrastructure on which Skild AI trains and deploys its model.[2]

04

Advertised Capabilities and Key Limitations to Keep in Mind

Capacity or Constraint What You Need to Know
One-time demonstration A video provides context for the execution. It does not guarantee consistent success on the first try.
No post-workout The model's weights are not adjusted for each new task presented in the experiments.
Time-consuming tasks The published demonstrations are up to ten minutes long and include several steps.
Adaptation According to Skild AI, S1 can handle certain object movements, substitutions, and runtime errors.
Evaluation The available figures come from Skild AI. No equivalent independent benchmarks have been published yet.
Availability S1 is available through select deployments and partnerships, not as a freely downloadable model.
05

What the results actually allow us to conclude

Skild AI compares a policy guided by video demonstrations to a VLA (Vision Language Action) model driven by a linguistic instruction. On unseen tasks lasting four to eight minutes, trained on a reported 100,000 hours of data, the company reports an average cumulative success rate per step of 66% for S1, compared to 9% for the linguistic baseline.[1]

This metric does not represent a 66% success rate for each mission. It aggregates the success rates of the various stages. Skild AI further notes that human intervention was used to reset the system after a failure so that all stages could be evaluated. This methodology provides information on the robot’s local progress, but it does not directly measure the proportion of tasks completed entirely without assistance.

The company also estimates that a real-world demonstration is equivalent to approximately 380 post-training sessions on the tasks studied. This equivalence is derived from an interpolation between several experimental data points. It should not be interpreted as a universal rule applicable to every task, piece of equipment, or environment.[1]

Published Results Reported value Read Carefully
Maximum duration displayed Up to 10 minutes A substantial amount of time is required for a robotics demonstration, focusing on a small number of selected scenarios.
New Tasks 66% per stage Internal measurement that includes human intervention following certain failures; it is not a rate of fully autonomous missions.
VLA Reference 9% per stage Internal comparison between two methods trained using the protocol defined by Skild AI.
Effectiveness of the Demonstration About 380 episodes An interpolated estimate that depends on the tasks, data, and equipment under study.
Tasks We've Already Covered About 96% Internal results on tasks similar to the pre-training distribution.

These results support a limited but important conclusion: in the Skild AI protocol, a video demonstration provides more information than a text-based instruction for certain long and novel tasks. They do not yet demonstrate that a general-purpose robot can learn any task from a video, nor that it can operate with the level of safety required in production.

As of September 22, 2026, no peer-reviewed publications or independent replications of all of S1’s results had been identified. The company’s videos and figures therefore constitute useful primary evidence, but are still insufficient to support a robust generalization.

Executive MBA in AI & Business Transformation

The MBA Redesigned for the Age of AI. For experienced executives who want to lead the transformation of their organizations. Paris, Nice, and Dubai.

12 months — Part-time At least 10 years of experience Early bird: 20,000 € Paris · Nice · Dubai
06

Why This Approach Matters to the Industry

In manufacturing and logistics, a significant portion of integration costs stems from adapting robots to products, workstations, and work sequences. If a video demonstration truly makes it possible to reconfigure a task in just a few minutes, production lines could become more flexible—particularly for small batches, frequent changes, or operations that are difficult to formalize through code.

The role of operators would also evolve. Their expertise in performing tasks could become a direct source of instructions for robotic systems. This would create new needs in terms of preparing demonstrations, validating trajectories, analyzing errors, and supervising deployments. Automation would therefore not necessarily eliminate human involvement, but would shift part of the work toward the transmission, control, and governance of actions.

Skild AI and NVIDIA cite applications in manufacturing, logistics, inspection, security, and food preparation. In particular, a partnership with Foxconn aims to apply Skild Brain to dual-arm robots for the assembly of NVIDIA Blackwell systems.[2] These announcements signal deployment, but publicly available data does not yet allow for an assessment of availability, ongoing reliability, cost per task, or incidents encountered in production.

07

A video alone is not enough to ensure widespread adoption

A demonstration covers only one possible trajectory. In a real workshop, objects change orientation, tools wear out, lighting varies, and people may enter the work area. Generalization requires that the model distinguish the objective from the exact sequence observed and be able to respond to situations not shown in the video.

Skild AI demonstrates tests in which S1 continues a task after an object has been moved or replaced. The company also states that the system can resume an action after certain failures. These behaviors are interesting, but they are still presented using selected examples. An industrial evaluation should document complete failure rates, near-misses, recovery conditions, differences between robots, and performance degradation over long periods.

Rapid adaptation also creates a validation issue. While any operator can teach a new task, the organization must decide who is authorized to do so, how the demonstration is verified, in what environment it can be performed, and how to revert to a previous version. Context-based learning potentially reduces programming time, but it does not eliminate the need for process qualification or the responsibility for deployment.

08

Safety, Liability, and Legal Framework

A robotic model operates in the physical world. An error does not merely produce incorrect text; it can damage an object, interrupt production, or injure someone. Mechanisms for force limiting, emergency shutdown, zone separation, speed monitoring, and human validation therefore remain essential, even when the action policy is generated by a foundation model.

In the European Union, the Artificial Intelligence Regulation is being phased in. An AI system integrated as a safety component into a product covered by European legislation may fall under the high-risk systems regime when the conditions of Article 6 are met. The obligations associated with systems covered by Article 6(1) are, in principle, to take effect on August 2, 2027.[4] Therefore, not all robots using S1 automatically become high-risk systems: classification depends on the function, the product, and the context of use.

The European Machinery Regulation, which is set to take effect primarily on January 20, 2027, also places greater emphasis on software and the safety functions of machinery.[5] For professional deployment, the assessment must address machine safety, cybersecurity, software change management, and requirements related to the AI system. This presentation is general in nature and does not constitute legal advice.

Responsibility must also be clearly allocated among the model provider, the robot manufacturer, the integrator, the employer, and the user. Execution logs, the model version, the demonstration video, safety parameters, and human interventions must be traceable when a decision or incident is analyzed.

09

Data, Labor, and Environmental Impact

Demonstration videos may capture employees, production areas, confidential processes, or protected items. Their collection must be limited to what is necessary, secured, and governed by rules regarding retention and access. When individuals are identifiable, the GDPR may apply. Organizations must also verify the rights to the videos and ensure the protection of their industrial know-how.

From a social perspective, observational learning can highlight operators’ expertise, but it can also lead to the recording of their actions without adequate recognition. Responsible implementation requires involving teams, documenting job changes, providing training in supervision, and ensuring that video does not become a tool for individual surveillance. Furthermore, the ergonomic quality of a human movement does not automatically mean it is safe or optimal for a robot.

Finally, Skild AI does not publish a comprehensive energy balance for the training or inference of S1. Simulation and the reuse of a general-purpose model may reduce certain physical data collection activities, but training on large volumes of data and the use of accelerated infrastructure come at a hardware and energy cost. Without comparable metrics, no net environmental benefit can be claimed.

10

What to Watch for Now

The next decisive step will not be another demonstration video, but the publication of reproducible protocols and independent results. In particular, it will be necessary to measure complete success without human intervention, robustness across different objects and robots, safety in the presence of people, downtime, computational cost, and the ability to explain failures.

We will also need to monitor the announced industrial deployments. A one-time demonstration and continuous production involve very different requirements. The most informative data will be that which describes how the system operates over several weeks, the number of human interventions, the necessary modifications, and the conditions under which the system must be shut down.

S1 thus represents a credible step forward in context-aware robotic learning, but it does not yet prove that a universal robot can learn instantly. Its true impact will depend on its ability to transform a promising demonstration into a reliable, verifiable, and accountable process.

Learn more

To understand S1 in the context of developments in robotics, embedded AI, and industrial sectors, be sure to check out these analyses from the aivancity blog.

Sources

[1] Skild AI, August 2026. Introducing S1: In-Context Learning for Robotics. Accessed September 22, 2026. View source

[2] NVIDIA, September 10, 2026. Skild AI Uses NVIDIA Physical AI to Teach Robots New Tasks From a Single Video. View source

[3] Robotics and Automation News, September 14, 2026. Skild AI Unveils S1 Robot Foundation Model That Learns Tasks From Video Demonstrations. Read the article

[4] European Commission. European Regulatory Framework on Artificial Intelligence. Accessed September 22, 2026. View the regulatory framework

[5] European Union, 2023. Regulation (EU) 2023/1230 on machinery. View the regulation

Exit mobile version