DeepSeek released V4.1-Flash on Thursday. It is a smaller model that the company says beats its own flagship on coding and agent tasks, at a lower price. The Hangzhou lab announced it in a thread on X and published the weights on Hugging Face under the MIT licence. That lets anyone download, modify and run it.
The launch comes with a notable retirement. From 04:00 UTC on 14 September, DeepSeek will send every request made to V4-Pro, its top model, to V4.1-Flash instead. Those requests will be billed at the cheaper Flash rates. That arrangement runs until a V4.1-Pro arrives. DeepSeek gave no date for it.
A big model that works like a small one
V4.1-Flash has 552 billion parameters in total. It uses only a small slice of them for each piece of text. DeepSeek calls the design a “causal encoder-decoder” and describes it as the smallest model in a new architecture family. It activates 8 billion parameters per token when reading input and 16 billion when writing output.
That split is aimed at AI agents. An agent that calls tools repeatedly spends much of its time reading fresh input, so cheaper reading means cheaper agents. The model handles up to 1 million tokens of context and understands images natively, according to its model card. DeepSeek says it pre-trained the model on 45 trillion tokens.












