Operating principle of Ouroboros. Conventional methods (left) detect visual changes caused by object and background motion and recompute patches accordingly, whereas Ouroboros (right) uses motion vectors from a video encoder to track identical visual information and selectively recompute only regions where actual changes occur. Red indicates recomputed patches, while green indicates reused patches. Credit: Seoul National University College of Engineering
A research team led by Kyunghan Lee, a professor in the Department of Electrical and Computer Engineering at Seoul National University College of Engineering, has developed an artificial intelligence system, "Ouroboros," that performs vision transformer-based video analysis 2.61 times faster using 13.0% of the computation required by conventional methods on edge devices.
The system reduces redundant computations in vision transformers by reusing visual information that repeatedly appears across consecutive video frames instead of recalculating it. In particular, it tracks objects and backgrounds as they move within a video, identifies identical visual information and selectively computes only regions with significant changes.
The team demonstrated that Ouroboros can significantly reduce computation, latency and energy consumption while preserving the accuracy of high-performance video AI models. Ouroboros is therefore expected to play a decisive role in accelerating the commercialization of real-time video AI in fields such as physical AI and augmented reality (AR), where power and computational resources are limited.









