Mastodon Skip to content Breaking US C-17 lands in Moscow with CIA chief Ratcliffe aboardUS C-17 lands in Moscow with CIA chief Ratcliffe aboardUS C-17 lands in Moscow with CIA chief Ratcliffe aboardUS C-17 lands in Moscow with CIA chief Ratcliffe aboardUS C-17 lands in Moscow with CIA chief Ratcliffe aboard
pulseofnations. Real News. Global Impact.
Subscribe
live markets
S&P 5007,677.28▲ 3.58%NASDAQ26,151.30▲ 4.71%DOW53,577.40▲ 3.14%GOLD4,716.00▲ 15.94%WTI80.40▼ 9.98%BRENT85.31▼ 11.85%EUR/USD1.1675▲ 2.62%USD/JPY158.95▼ 2.98%DXY98.91▼ 2.53%BTC$78,878▼ 2.35%ETH$2,460▼ 2.13%SOL$97.08▼ 4.30%

DeepSeek Adds Vision to V4 Flash, Matches Opus 4.8 on Benchmarks

DeepSeek releases V4-Flash-Vision-Exp, a multimodal model that adds image understanding to its cheap text workhorse while matching Anthropic Opus 4.8 on visual agent benchmarks.

Partner Surfshark VPN

DeepSeek has released V4-Flash-Vision-Exp, an experimental multimodal model that adds image and screenshot understanding to its V4-Flash text workhorse, claiming performance nearly on par with Anthropic’s Opus 4.8 on visual agent benchmarks. The model, available via DeepSeek’s API under the ID deepseek-v4-flash-vision-exp, was launched on August 21 and ships with DeepSeek Harness 0.1.1 out of the box.

What Changed

The new model builds on V4-Flash’s existing text and agent capabilities by adding visual understanding. According to DeepSeek’s own benchmark data, V4-Flash-Vision-Exp matches V4-Flash on every text benchmark – within a point or two on six of seven evaluations – confirming the company didn’t trade text quality for vision ability.

On the multimodal side, the results are more striking. V4-Flash-Vision-Exp actually beats Opus 4.8 on two of four multimodal agent benchmarks: Agents’ Last Exam (27.3 vs 25.7) and ZeroBench Pass@5 (35.0 vs 34.0). It trails on ApexBench (36.5 vs 39.4) and Chartography (64.3 vs 65.0). Crucially, these aren’t simple image-captioning tests – they are agentic tasks that require reading screenshots, charts, or diagrams as part of a longer tool-use loop.

The Price Story

The most interesting detail for developers may be the pricing. DeepSeek is not charging a separate rate for image tokens. Each image consumes up to 384 input tokens and is billed at V4-Flash’s existing per-million-token rate – off-peak input at $0.007/M on cache hit or $0.22/M on cache miss. That effectively makes vision a free upgrade for anyone already running V4-Flash for text-agent work, eliminating the need to route screenshot-reading steps to a separate vision model.

The model supports JPEG, PNG, GIF and WebP formats, accepts up to 600 images per request, and works across Chat Completions, Anthropic-compatible Messages, and OpenAI-compatible Responses API endpoints. The experimental designation signals DeepSeek is gathering feedback before committing to a production release.

The release arrives amid escalating competition in the multimodal AI space, where vision capabilities are drifting from a flagship perk to a baseline expectation. Google’s Gemini 3.7 Flash, launched August 13, is natively multimodal with text, images, video and audio support. Anthropic and OpenAI have both shipped vision-capable models throughout 2025 and 2026.

The timing is also notable given DeepSeek’s broader trajectory this month. The company raised API prices 50 to 1,100 percent across the V4 series in mid-August, and speculation about an IPO at roughly $86 billion continues to swirl. The vision variant extends V4-Flash from a text-only budget model into a multimodal agent platform at the same price point, potentially widening its appeal among developers building screenshot-reading and UI automation workflows.

SourcesDeepSeek official announcement on X, August 21; explainx.ai benchmark analysis, August 21; emergent.sh, August 22; iWeaver.ai technical breakdown, August 22
React to this dispatch
Share this dispatch X WhatsApp Bluesky Report an error
Written by

Founder and editor of Pulse of Nations, an independent wire service covering war, geopolitics, markets and technology.

discussion

Leave a Reply

Next dispatch NVIDIA Groq 3 LPX Enters Production at Record Speed Read →