DeepSeek V4.1 Flash replaces V4 Pro API traffic after September 14
DeepSeek says V4.1 Flash beats V4 Pro on cost and several agent tests, but its own results show Pro ahead on two reasoning measures.
By Colin Brandt · Enterprise Reporter
· 3 min read
DeepSeek V4.1 Flash was released September 10 as the Chinese AI company’s newest API model, and DeepSeek says it will redirect all V4 Pro endpoint traffic to the cheaper model at 04:00 UTC on September 14. The change makes the launch more than a new model option: customers using the existing V4 Pro endpoint will receive V4.1 Flash, billed at Flash rates, until DeepSeek releases V4.1 Pro.
DeepSeek said its testing found V4.1 Flash ahead of V4 Pro on performance, cost, speed and total runtime. That is a company-reported assessment, rather than a settled independent conclusion. Its published benchmark results favor the new Flash model on several agentic, coding and cybersecurity measures, while also showing V4 Pro ahead on two knowledge and reasoning tests.
When will DeepSeek switch V4 Pro traffic to V4.1 Flash?
The switchover is scheduled for September 14 at 12:00 Beijing Time, or 04:00 UTC. Requests made to the deepseek-v4-pro endpoint will then be served by V4.1 Flash and charged using its rates, according to DeepSeek’s API changelog. The arrangement is temporary, lasting until V4.1 Pro becomes available.
DeepSeek has already retired V4 Flash and V4 Flash Vision Exp. Applications still sending requests under the legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names are being temporarily routed to V4.1 Flash for compatibility. Compatibility routing keeps those names working, but does not itself demonstrate output parity. Teams with production workflows tied to either retired Flash models or V4 Pro should test outputs, latency and spend around the transition.
DeepSeek’s benchmarks show a mixed comparison
On matching results disclosed by DeepSeek, V4.1 Flash scored 90.6 on Terminal-Bench 2.1, versus 87.9 for V4 Pro; 74.2 versus 62.7 on DeepSWE v1.1; and 88.1 versus 83.3 on CyberGym. Those results support the company’s claim of gains on the coding, agent and security-oriented tests it selected.
The vendor’s comparisons are not a universal win. Its published table puts V4 Pro ahead on GPQA Diamond, 92.4 to 90.9, and on the cited pure-text subset of Humanity’s Last Exam, 42.7 to 39.1. Benchmark labels and test configurations can differ across releases, so operators should not treat the aggregate claim as a substitute for task-specific evaluation.
What changed in V4.1 Flash?
DeepSeek describes V4.1 Flash as the smallest model in a new architecture family, with native multimodal visual understanding. The mixture-of-experts model has 552 billion parameters and uses what the company calls a Causal Encoder-Decoder design. It activates 8 billion parameters when processing input and 16 billion when generating output.
The company also attributes lower serving costs to a smaller key-value cache, the stored context state used during inference. DeepSeek says the new cache requires one-quarter of the prior generation’s HBM and one-eighth of its SSD storage. The model is available through the DeepSeek API under the deepseek-flash name. New API pricing took effect September 10, with off-peak rates set at half of peak rates; DeepSeek did not provide text pricing figures in its launch announcement.
This story draws on original reporting from SiliconANGLE.