DeepSeek V4.1 Flash
DeepSeek · released 10 Sep 2026 · open weights
- Newest release
- Open weights, self-hostable
552B MoE that activates only 8B parameters on input and 16B on output, with native visual understanding; DeepSeek's own benchmarks put it ahead of V4 Pro on quality, cost and speed at a quarter of the price.
Context window
1M
1,000,000 tokens
Input price
$0.30
per 1M tokens
Output price
$1.20
per 1M tokens
Blended
$0.52
3:1 in:out mix
Weights
Open
self-hostable
Released
10 Sep 2026
DeepSeek
Long-context / tier pricing — deepseek-flash, peak rate, cache-miss input. Off-peak is exactly half ($0.15/$0.60) and covers every hour outside 01:00-04:00 and 06:00-10:00 UTC on weekdays, so schedule batch work off peak. Cache hits are $0.006 peak. Max output 384K. Listed prices are the base rate for prompts up to ~200K tokens.
Open-weights API prices are representative third-party-hosted rates; self-hosting costs only your own hardware.
Where it ranks
Of 28 models tracked.
Cost to run
#4
of 28
Input price
#4
of 28
Context window
#6(tied)
of 28
Compare with
The nearest models by blended cost to run.
Put it to work
Specs from BitByteCore's hand-verified Frontier Ledger (v2026-09-05), this row verified 10 Sep 2026 against DeepSeek's official pricing ↗. Model pricing moves fast, so confirm with DeepSeek before you rely on a figure.