Learning factory tasks from a single video enables robots to shift from reprogramming to instant operation with the S1 model
Manufacturing facilities, warehouses, and supply chains face a persistent challenge of shifting work requirements, changing production routes, and successive new products, which has historically imposed high costs and long downtime for reprogramming industrial robots and training them from scratch. In a step aimed at breaking this complex operational cycle, Scaled AI launched its new foundational model “S1” for physical AI, which allows robots to absorb long, complex tasks and execute them after watching a single instructional video, relying on contextual learning without modifying model weights and without needing subsequent training stages for each task individually.
The radical shift here moves robots from pre-programmed logic to direct experiential understanding.Deback Patak, co-founder and CEO, explains that the pipeline converts the video recorded by a human operator into a practical prompt that interprets intentions, object movements, and step sequences, then translates them into precise physical motions performed directly by the robot. In applied tests, the model demonstrated the ability to complete unfamiliar tasks lasting up to ten minutes, such as moving seedlings to trays, making pastries, preparing drip coffee, and assembling part kits, processes that involve dozens of manual steps and require the robot to autonomously adjust motion and regain balance when objects shift or unexpected disturbances occur.
Performance data reveal an exceptional efficiency gap in automation settings, with the transition from video capture to full autonomous operation on hardware taking only eleven minutes in a seedling-planting trial. The robot also achieved a success rate of roughly 66 % at each step of multi-stage tasks it had not been previously trained on, versus about 9 % for comparable AI systems, delivering more than seven-fold superiority. The company estimates that showing a single short video to the model is equivalent to providing it with approximately 380 practical training examples, saving operation teams between 50 and 100 hours of labor-intensive manual data collection for each production-line change.
The model has moved beyond laboratory trials, entering precise assembly lines in collaboration with NVIDIA and Foxconn.The three parties are operating the Scaled system with dual robotic arms to assemble advanced NVIDIA Blackwell systems, where the robot installs electrical bus bars, fastening parts, and secures sixteen screws with extreme precision while adapting to any changes in part location. The system’s construction relies on NVIDIA computing infrastructure, leveraging Cosmos models to convert video into structured data, and the OmniVerse and Izac Lab libraries to simulate contact, collision, and pressure forces via the Newton physics engine, culminating in the TensorRT suite to accelerate real-world motion inference. This development coincides with the company reaching an annual revenue rate of $100 million only ten months after its first commercial deployment, and signing more than 60 distribution partnerships across manufacturing, logistics, inspection, security, and food preparation sectors.
This shift in automation flexibility carries direct practical implications for industrial and logistics facilities in the Gulf, Egypt, and the Levant. Traditionally, the cost of introducing robots in regional packaging, assembly, and storage plants has been concentrated in reliance on external system-integration firms to reprogram code and test paths for every product change. When the cost of aligning a machine drops to the level of recording a single video, medium-sized facilities can reduce capital expenditures and avoid weeks-long production-line shutdowns. The required skill set for local engineering staff also changes: the priority moves from writing precise code for each arm’s trajectory to visual guidance, simulation environment management, and verification of inference efficiency, enabling broader robot deployment in flexible workspaces without costly software investments.
The factory operating equation changes when the human operator becomes an instant trainer for the machine via a regular camera.Your practical step as an operations or supply-chain manager is to reassess the viability of automating variable tasks previously excluded due to high programming costs, as contextual learning opens the door to deploying robots on multiple tasks within a few minutes.