DeepSeek published the official public beta of its V4-Flash API on July 31 under the designation V4-Flash-0731, marking the end of a preview period that began in April. The model, a 284-billion-parameter mixture-of-experts system with 13 billion active parameters per token, now carries a major upgrade: it beats the flagship 1.6-trillion-parameter V4-Pro on nine agent benchmarks.
The release is notable not for a bigger model or new architecture, but for a retrained one. DeepSeek invested additional compute into refining V4-Flash’s training process, yielding significant gains in agentic reasoning and coding tasks while maintaining the same efficient inference profile that made the model attractive to developers in the first place.
The model supports a one-million-token context window and operates in both thinking and non-thinking modes. Under an MIT license, the weights remain fully open, allowing developers to self-host and modify the model without restrictions. The API pricing positions it as one of the most cost-effective frontier-class models available, at roughly $0.14 per million tokens.
DeepSeek had originally targeted a mid-July release for the official V4-Flash, but the additional retraining pushed the timeline back by two weeks. The separate V4-Pro iteration remains available but has not received the same retraining treatment, creating an unusual situation where the smaller model outperforms the larger one on several key metrics.
The release intensifies China’s AI model competition, with DeepSeek going head-to-head against Alibaba’s Qwen series and ByteDance’s Doubao models. The company’s ability to deliver high-performance models at dramatically lower price points than US-based competitors continues to reshape expectations around AI inference costs.
Developers can access V4-Flash through both OpenAI-compatible and Anthropic-compatible API interfaces, with the base URL unchanged from previous versions. The model card on Hugging Face explicitly states this release supersedes the April preview, though the preview endpoint will remain operational through the end of August.
Sources: MarkTechPost, TechTimes, DeepSeek API Docs
Author: Technology Desk
discussion