VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

DeepSeek Opens One-Day Beta of V4.1 Flash Multimodal Model

DeepSeek launched a limited beta of V4.1 Flash, an interim multimodal model with a new architecture, available via API until September 10 at V4 Flash…

DeepSeek has begun a limited-time beta of V4.1 Flash, an interim model that natively supports multimodal capabilities and uses a new architecture1. The beta window is narrow: the model is scheduled to go offline on September 10.

DeepSeek said V4.1 Flash offers stronger performance and faster generation at a lower cost relative to its predecessor, though the company has not presented the release as a formal launch. A separate Chinese-language report described the model as delivering "stronger visual intelligence, extreme speed, and cost efficiency"2.

Developers can access V4.1 Flash through DeepSeek's existing API by selecting the endpoint deepseek-v4.1-flash-expires-on-0910. Beta pricing matches V4 Flash rates, and each account supports up to 20 concurrent requests.

The beta arrives during a period of active infrastructure and model development at DeepSeek. The company recently ordered more than 160,000 Huawei Ascend 950DT accelerators for a data center under construction in Inner Mongolia[2], and NVIDIA's Dynamo project published experimental inference recipes for DeepSeek V4-Pro-0813 on GB200 and H200 hardware[3]. Red Hat also recently detailed FastMTP heads for vLLM speculative decoding, an optimization technique tied to the multi-token prediction training objective used by DeepSeek models[1].

ANALYSIS Labeling V4.1 Flash as an interim model with a single-day beta window and an expiring API endpoint suggests DeepSeek is using the window to gather external feedback or stress-test the new architecture before a broader release. Native multimodal support marks a capability expansion for the Flash line, which has previously been positioned on speed and cost rather than visual reasoning.