VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

DeepSeek V4 Pro Hits General Availability With 1M Context Window

DeepSeek's V4 Pro model reaches general availability with a 1M-token context window, 384K output, Responses API support, and tiered pricing against Flash.

Vector Wire — AI-assisted editorial illustration

DeepSeek has pushed its V4 Pro model to general availability, versioned as DeepSeek-V4-Pro-0813, upgrading the April preview release to a full production offering1,6. The flagship model ships with a 1M-token context window and a 384K-token maximum output2, positioning it as the company's most capable model for long-document ingestion and extended generation tasks.

DeepSeek's official website described the release as delivering "significantly enhanced agent capabilities". The model defaults to a thinking mode while also exposing a non-thinking endpoint for latency-sensitive calls. It supports structured JSON output, tool calls, the Responses API, and Anthropic-API compatibility, plus beta support for conversation-prefix continuation and fill-in-the-middle (FIM) completion, with FIM restricted to the non-thinking mode.

At the wire level, the service exposes both OpenAI-compatible and Anthropic-compatible endpoints, allowing teams already running on either standard to integrate without re-plumbing their stack.

Pricing reflects a deliberate tier split within the V4 family. Pro is listed at 0.025 yuan per million tokens for cache hits, 3 yuan for cache misses on input, and 6 yuan on output. In dollar terms, DeepSeek's API documentation lists the Pro tier at $0.003625 per million input tokens for cache hits, $0.435 for cache misses, and $0.87 per million output tokens. The Information reported comparable figures of $0.44 per million input tokens and $0.87 per million output tokens5. Flash, by contrast, is priced at 0.02, 1, and 2 yuan respectively, putting Pro at roughly three times the cost of Flash on cache-miss input and output.

Concurrency limits reinforce the product segmentation: Flash supports 2,500 concurrent requests, while Pro is capped at 500. ANALYSIS The concurrency and pricing gap suggests DeepSeek is steering high-throughput, cost-sensitive workloads to Flash while reserving Pro for heavier reasoning and agent tasks that benefit from the larger output budget.

The model is also available on OpenRouter3.

Early reception has been mixed. The Information framed the launch as a challenge to Moonshot AI's Kimi K3, reporting that V4 Pro rivals Kimi K3 on some benchmarks at lower prices4. The South China Morning Post, however, reported that some developers were "underwhelmed by its overall capabilities and disappointed in its pricing," while researchers found the model performed well in niche areas such as cybersecurity.

ANALYSIS The split reception — competitive on select benchmarks yet underwhelming to some practitioners — mirrors the tension inherent in pricing a model above Flash while competing against frontier-tier rivals. The cybersecurity strength noted by researchers may carve out a differentiated use case even if broad benchmark performance does not uniformly lead the field.

The release lands on a busy day for model launches, with social-media posts noting simultaneous availability of Grok 4.6 and Qwen 3.8 models, though those claims remain unverified8.

CORRECTIONS: none for this article · this piece updates automatically as the story develops · corrections policy & trail →