Skip to content

DeepSeek V4.1 Flash

DeepSeek · released 10 Sep 2026 · open weights

552B MoE that activates only 8B parameters on input and 16B on output, with native visual understanding; DeepSeek's own benchmarks put it ahead of V4 Pro on quality, cost and speed at a quarter of the price.

Context window

1M

1,000,000 tokens

Input price

$0.30

per 1M tokens

Output price

$1.20

per 1M tokens

Blended

$0.52

3:1 in:out mix

Weights

Open

self-hostable

Released

10 Sep 2026

DeepSeek

Long-context / tier pricingdeepseek-flash, peak rate, cache-miss input. Off-peak is exactly half ($0.15/$0.60) and covers every hour outside 01:00-04:00 and 06:00-10:00 UTC on weekdays, so schedule batch work off peak. Cache hits are $0.006 peak. Max output 384K. Listed prices are the base rate for prompts up to ~200K tokens.

Open-weights API prices are representative third-party-hosted rates; self-hosting costs only your own hardware.

Where it ranks

Of 28 models tracked.

Cost to run

#4

of 28

Input price

#4

of 28

Context window

#6(tied)

of 28

See the full Frontier Index →

Compare with

The nearest models by blended cost to run.

Put it to work

Specs from BitByteCore's hand-verified Frontier Ledger (v2026-09-05), this row verified 10 Sep 2026 against DeepSeek's official pricing ↗. Model pricing moves fast, so confirm with DeepSeek before you rely on a figure.