DeepSeek has launched DeepSeek-V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native vision, a 1 million-token context window and pricing built around low cached-input costs.. During off-peak hours, DeepSeek prices the model at $0.003 per million input tokens on a cache hit, $0.15 per million on a cache miss and $0.60 per million output tokens.
Peak rates are double those figures, and DeepSeek defines peak hours as Monday through Friday from 01:00 to 04:00 UTC and 06:00 to 10:00 UTC.. DeepSeek said cache-hit charges can account for a significant share of agent costs when systems repeatedly read the same repository, tool definitions, system instructions or conversation history..
Using published prices cited in the source material, OpenAI lists GPT-5.6 Sol at $4 per million regular input tokens, $0.40 for cached input and $20 for output. Anthropic charges $5 for standard input, $0.50 for Claude Opus 5 cache hits and $25 for output, while Moonshot AI’s Kimi K3 costs $3 for cache-miss input, $0.30 for cache-hit input and $15 for output..
In a simplified example of a 500,000-token reusable prefix cached across 100 requests, those 50 million cached input tokens would cost about $0.15 on V4.1-Flash off-peak. The comparable cache-read bill would be about $15 on Kimi K3, $20 on GPT-5.6 Sol and $25 on Claude Opus 5, excluding cache-write charges, fresh context and output..
At off-peak rates, V4.1-Flash totals $0.75 per 1 million tokens based on $0.15 input and $0.60 output pricing. In the comparison table in the source material, that places it behind only Meta’s Contributor-tier Muse Spark at $0.30 total and Xiaomi’s MiMo-V2.5 Flash at $0.40, while its peak-hour total of $1.50 ties with MiniMax-M3 and LongCat-2.0 promotional pricing..
DeepSeek’s documentation says legacy V4 Flash requests are now served by V4.1-Flash. It also says V4.1-Flash has “comprehensively surpassed V4 Pro in performance, cost, speed, and total time,” and that V4 Pro will route to V4.1-Flash after Sept. 14, 2026 until a future V4.1 Pro release..
The architecture splits 40 Transformer layers into a 20-layer causal encoder and a 20-layer decoder. DeepSeek said the design activates 8 billion parameters per token during input processing and 16 billion during output generation.. DeepSeek said the model combines that design with Compressed Sparse Attention 2, hierarchical sparse indexing and FP4 KV caching.
The company said those techniques reduce the global KV cache to 890 bytes per token, about one-quarter the size of V4-Flash’s, while persistent cache storage falls to roughly one-eighth.. The model is larger than its predecessor in total weights. DeepSeek’s previous V4-Flash used a 284 billion-parameter backbone with 13 billion active parameters, while V4.1-Flash raises the backbone to 552 billion parameters, and the technical report separately lists 196 billion parameters in sparsely accessed Engram conditional-memory modules..
DeepSeek said the new architecture creates untested robustness limits. The company said sparse-selection errors and approximate state reconstruction could degrade capability in edge cases, particularly in sparse retrieval over very long contexts and cache-resumption boundaries, and said it plans additional stress testing..
On benchmarks run by DeepSeek, V4.1-Flash scored 74.2 on DeepSWE v1.1, compared with 74.0 for Claude Opus 5 and 73.0 for GPT-5.6 Sol. DeepSeek also reported scores of 88.1 on CyberGym and 54.8 on AutomationBench, while saying higher reasoning effort improves results but can consume about 2.5 times as many output tokens when raised from 25 to 100..
In VentureBeat Pulse Research’s July 2026 survey of 170 enterprises with more than 100 employees, 47% said they rigorously track AI compute cost and return on investment, while 53% said they do not. The survey also found 12% had not addressed inference-memory limits such as KV-cache capacity and 7% were not aware of the constraint..
Reuters reported this week that DeepSeek has tapped CITIC Securities as it prepares for a potential listing on Shanghai’s STAR Market. Reuters also reported that a current fundraising could value the company at as much as 500 billion yuan, or about $75 billion.. Featured image credit.
Tags: deepseekFeatured
