Deepseek releases V4.1-Flash, a multimodal model with 552 billion parameters that cuts KV cache memory to a quarter of its predecessor. On the DeepSWE coding benchmark, it narrowly beats Opus 5 and GPT-5.6 Sol, even though only 16 billion parameters are active per token. The model ships under the MIT license and targets much cheaper AI agents.

DeepSeek has begun a limited-time beta of V4.1 Flash, an interim model that uses a new architecture and natively supports multimodal capabilities.

DeepSeek V4.1 Flash: The Native Multimodal Model That's Breaking Speed Records ...