How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.

Shanghai-based company says M3 can process data five times faster than its predecessor, while also slashing inference costs.

Chinese start-up MiniMax has launched M3, an AI model with a redesigned architecture that reduces computational needs by up to 95%, enhancing efficiency and response speeds.