Back to Articles

1. From Surgical World Model to Interactive Simulator 2. Distilling Cosmos-H-Surgical-Simulator for Real Time 2.1. A Surgical Teacher 2.2. Causal Warmup 2.3. Self-Forcing Distillation 3. FlashDreams: The Real-Time Inference Engine 4. Adapting to Your Own Data 5. What Is Next: Toward Closed-Loop Surgical Physical AI 6. Get Started Today Surgical robotics is moving quickly from teleoperation toward increasingly capable vision-language-action policies. But evaluating and training these systems remains difficult. Physical robotic platforms are expensive to operate, experiments are slow to reproduce, and failures can damage instruments or biological material. Conventional simulators provide a safer alternative, but surgical scenes are exceptionally difficult to model: deformable tissue, fine instrument interactions, specular surfaces, sutures, needles, smoke, and occlusions all matter.

World foundation models offer a different path. Instead of manually authoring every object and physical interaction, they learn visual dynamics directly from synchronized video and robot kinematics. NVIDIA's Cosmos-H-Surgical-Simulator demonstrated this approach by generating future surgical video from an initial scene and a sequence of robot actions. It enabled faster-than-physical evaluation and synthetic data generation across the Open-H-Embodiment ecosystem.