DeepSeek-V4.1-Flash debuts with $0.003 off-peak cached input

Posted under: AI technologies
Date: 2026-09-11
DeepSeek-V4.1-Flash debuts with $0.003 off-peak cached input | Justo Global

DeepSeek has launched DeepSeek-V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native vision and a 1-million-token context window. During off-peak hours the model costs $0.003 per million input tokens on a cache hit, $0.15 on a cache miss and $0.60 per million output tokens. Peak rates are double those figures. Peak windows run Monday to Friday from 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak. The low cached-input price matters for agents that repeatedly read the same repository, tool definitions or conversation history. A simplified example of 50 million cached input tokens would cost about $0.15 off-peak on V4.1-Flash versus $20 on GPT-5.6 Sol and $25 on Claude Opus 5 at their published cache-read rates. The model uses a Causal Encoder-Decoder architecture that activates 8 billion parameters during prefill and 16 billion during generation. DeepSeek says the global KV cache has been reduced to 890 bytes per token.

Read more at: venturebeat.com

Related videos

DeepSeek-V4.1-Flash debuts with $0.003 off-peak cached input

Posted under: AI technologies
Date: 2026-09-11
DeepSeek-V4.1-Flash debuts with $0.003 off-peak cached input | Justo Global

DeepSeek has launched DeepSeek-V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native vision and a 1-million-token context window. During off-peak hours the model costs $0.003 per million input tokens on a cache hit, $0.15 on a cache miss and $0.60 per million output tokens. Peak rates are double those figures. Peak windows run Monday to Friday from 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak. The low cached-input price matters for agents that repeatedly read the same repository, tool definitions or conversation history. A simplified example of 50 million cached input tokens would cost about $0.15 off-peak on V4.1-Flash versus $20 on GPT-5.6 Sol and $25 on Claude Opus 5 at their published cache-read rates. The model uses a Causal Encoder-Decoder architecture that activates 8 billion parameters during prefill and 16 billion during generation. DeepSeek says the global KV cache has been reduced to 890 bytes per token.

Read more at: venturebeat.com
Open-source: The power of collective information

Open-source: The power of collective information

Open-source: The power of collective infor...

Elevate Your Sales Using Managed Services - Don't Miss Out!

Elevate Your Sales Using Managed Services - Don't Miss Out!

Elevate Your Sales Using Managed Services ...

How CRM Transforms Customer Relationships? #crm #technology #technews #business #businessautomation

How CRM Transforms Customer Relationships? #crm #technology ...

How CRM Transforms Customer Relationships?...