A company called Catnip has released what it’s calling MaineCoon, a real-time audio-visual AI model boasting 22 billion parameters that can reportedly create live character streams, complete with synced speech and motion, from nothing more than text prompts.

What MaineCoon claims to do

The core pitch is straightforward: type a text prompt, and MaineCoon generates a live video stream of a character speaking and moving in real time. Not a pre-rendered clip. Not a deepfake stitched together from existing footage. A continuous, live stream where the audio and visual components are generated and synchronized simultaneously.

The verification problem

As of now, independent verification of MaineCoon’s capabilities remains elusive. No major AI research outlets, crypto media platforms, or technology publications have published corroborating coverage of the model or its release. No technical paper, model card, or benchmark results have surfaced publicly.