TL;DR

Frame-only video search can't find "where she explains pricing." Transcript-only search can't find "the architecture diagram slide." We'll build both indexes, key them to the same timeline, and fuse at query time so results are timestamps, not video IDs. Shot detection, SigLIP-2 embeddings, Whisper chunks, pgvector.

What we're building

A search endpoint where "where does she explain the pricing model" returns video_id=42, t=247.3s and you can seek straight there.

Two indexes over one timeline: