A user opened an issue on my open-source video tool last week that named a gap I had been shipping around for months.

He runs lectures through crv so an LLM can read them. One 22-minute lecture: 1,377 candidate frames extracted, 60 kept after dedup and --max-frames thinning. The 60 frames come out in the right order. That is all they come out with.

His complaint, in one line: the LLM can describe the slide, but it cannot tell you when the slide was on screen.

That breaks more than it sounds like:

you cannot cite visual evidence with a timestamp