TL;DRHundreds of millions of cameras are already deployed, but most footage still requires a human to know what to look for. Lumana, founded by ex-Intel computer vision leaders, processes over a billion images daily across 50,000+ cameras using its VIA-1 model, which learns what is normal for each individual camera and flags deviations. The company filters locally before sending anything to the cloud, following a principle of “filter before you spend.” Video may be physical AI’s natural starting point because the infrastructure already exists.
A camera overlooking a loading dock might record twelve hours of trucks arriving, workers moving through the site, and boxes leaving the building. Most days, nobody has a reason to watch any of it. The footage only becomes useful when a package goes missing, an accident happens or somebody needs to work backwards from an event and find out what happened.
That has been one of the strange limitations of video surveillance for years. Cameras became digital long ago, but the footage they produce still depends heavily on a person knowing what to look for and where to find it. AI is beginning to make more of that footage understandable and searchable while events are still unfolding.






