DeepSeek releases experimental multimodal V4-Flash-Vision-Exp
DeepSeek
DeepSeek shipped V4-Flash-Vision-Exp on its API platform as its first multimodal entry, matching V4-Flash on text tasks (agents, reasoning, world knowledge) while adding image understanding billed at up to 384 tokens per image. The vision model lands close to Opus-4.8 on DeepSeek's multimodal-agent benchmark table and ships with a free Files API for image reuse and DeepSeek Harness 0.1.1 day-one support.
Why it matters
Marks DeepSeek's move beyond text into the multimodal race at Flash-tier pricing, with day-one coverage in Chinese financial press (Caixin) framing it as the start of multimodal competition alongside Moonshot and Qwen.
Importance: 3/5
Frontier Chinese lab first-multimodal release with 4 independent sources (2 official + 2 primary-media)