DeepSeek launches V4.1-Flash, a 552B-parameter MoE model activating just 8-16B parameters with a 1M token context window, rivaling GPT-5.6

DeepSeek has begun a limited-time beta of V4.1 Flash, an interim model that uses a new architecture and natively supports multimodal capabilities.

DeepSeek V4.1 Flash: The Native Multimodal Model That's Breaking Speed Records ...