Author(s): Chew Loong Nian - AI ENGINEER Originally published on Towards AI. MiniMax M3 Decodes 1M Tokens 15x Faster — and It Shouldn't Be This Cheap O ...

Shanghai-based company says M3 can process data five times faster than its predecessor, while also slashing inference costs.

Chinese start-up MiniMax has launched M3, an AI model with a redesigned architecture that reduces computational needs by up to 95%, enhancing efficiency and response speeds.