- Sources: primary, pricing, discussion
- Summary: DeepSeek's API pricing page lists DeepSeek-V4.1-Flash under the model id
deepseek-flash, with a 1M context window, 384K maximum output, and vision support that deepseek-v4-pro does not have. Off-peak pricing is $0.15 per 1M input tokens on a cache miss and $0.60 per 1M output, against $0.66 and $1.98 for deepseek-v4-pro, and the concurrency limit is 2500 against 500. DeepSeek states that V4.1 Flash has comprehensively surpassed V4 Pro in performance, cost, speed, and total time, and that from 12:00 Beijing time on 2026-09-14 deepseek-v4-pro requests route to V4.1 Flash at the Flash price until a V4.1 Pro release. - Why it matters: Teams pinned to
deepseek-v4-pro get moved onto a different model at a different price on 2026-09-14 without changing the model id they send. - Follow-up: The V4 Pro retirement on 2026-09-14, and whether a V4.1 Pro release follows.
send feedback on this story