A user opened an issue on my open-source video tool last week that named a gap I had been shipping around for months.
He runs lectures through crv so an LLM can read them. One 22-minute lecture: 1,377 candidate frames extracted, 60 kept after dedup and --max-frames thinning. The 60 frames come out in the right order. That is all they come out with.
His complaint, in one line: the LLM can describe the slide, but it cannot tell you when the slide was on screen.
That breaks more than it sounds like:
you cannot cite visual evidence with a timestamp






