DeepSeek shipped V4-Flash-0731 last week — same 284B parameter architecture as the preview, same 13B activated parameters per token, MIT licensed, open weights on HuggingFace. No architecture changes. No bigger model.
It now outperforms V4-Pro-Preview on several agent benchmarks.
"We've massively upgraded its Agent capabilities — benchmark scores are now far surpassing the V4-Pro-Preview."
That's what makes this release interesting. Not the model. The method.
What actually changed







