Chinese AI company Deepseek has released V4-Flash-Vision-Exp, an experimental multimodal model that adds image understanding to its text capabilities. On Deepseek's own benchmarks, the model nearly matches Opus 4.8 on agent tasks.

Deepseek-V4-Flash-Vision-Exp extends Deepseek-V4-Flash with image processing while keeping the base model's text performance in reasoning and world knowledge, Deepseek says. On the company's internal multimodal agent benchmarks, the vision variant scores close to Opus 4.8.

Deepseek is targeting visual agent workflows

Deepseek is positioning the model for agent-based applications. It's designed to work with different agent frameworks and combine visual understanding with tool use. In practice, it can describe images, extract text from screenshots, and analyze diagrams. It handles JPEG, PNG, GIF, and WebP, and determines the format from actual file content rather than the filename or declared MIME type, per the API docs.

With vision capabilities added, the new Deepseek model approaches or sometimes beats Opus 4.8 on agent benchmarks. | Image: Deepseek