DeepSeek has launched DeepSeek-V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native vision and a 1-million-token context window. During off-peak hours the model costs $0.003 per million input tokens on a cache hit, $0.15 on a cache miss and $0.60 per million output tokens. Peak rates are double those figures. Peak windows run Monday to Friday from 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak. The low cached-input price matters for agents that repeatedly read the same repository, tool definitions or conversation history. A simplified example of 50 million cached input tokens would cost about $0.15 off-peak on V4.1-Flash versus $20 on GPT-5.6 Sol and $25 on Claude Opus 5 at their published cache-read rates. The model uses a Causal Encoder-Decoder architecture that activates 8 billion parameters during prefill and 16 billion during generation. DeepSeek says the global KV cache has been reduced to 890 bytes per token.
DeepSeek has launched DeepSeek-V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native vision and a 1-million-token context window. During off-peak hours the model costs $0.003 per million input tokens on a cache hit, $0.15 on a cache miss and $0.60 per million output tokens. Peak rates are double those figures. Peak windows run Monday to Friday from 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak. The low cached-input price matters for agents that repeatedly read the same repository, tool definitions or conversation history. A simplified example of 50 million cached input tokens would cost about $0.15 off-peak on V4.1-Flash versus $20 on GPT-5.6 Sol and $25 on Claude Opus 5 at their published cache-read rates. The model uses a Causal Encoder-Decoder architecture that activates 8 billion parameters during prefill and 16 billion during generation. DeepSeek says the global KV cache has been reduced to 890 bytes per token.