- Sources: DeepSeek vision guide, DeepSeek API change log, HN discussion
- Summary: The change log dates the release to 2026-08-21, the model carries an
exp suffix, and images are accepted only in user messages and only by this model, through the OpenAI-compatible endpoint, the Anthropic-compatible /messages endpoint, and the Responses API. Every image is resized to roughly the pixel count of an 800 by 800 image before inference, which puts a hard ceiling of 384 tokens per image, so a 2000 by 2000 and a 5000 by 5000 image cost the same, against published limits of 600 images per request, a 48 MiB request body, 32 MiB per image inline or by URL versus 64 MiB through the Files API, and a maximum of 8192 px per side that drops to 4096 px once a request carries 15 or more images. Every benchmark figure is DeepSeek's own, including Terminal Bench 2.1 at 83.9 and ZeroBench Pass@5 at 35.0 published with a stated method for the code-agent text tasks, and DeepSeek's statement that the model's multimodal agent capability comes close to Opus-4.8 is a vendor comparison with no independent reproduction located. - Why it matters: A fixed 384-token ceiling per image makes image cost predictable and independent of resolution, so sending a larger image buys no additional detail at inference.
send feedback on this story