VECTOR WIREAI INTELLIGENCE
NVDA$1,847+3.2%MSFT$512+1.1%GOOGL$199-0.4%META$728+2.7%AMD$184-1.2%TSM$212+0.6%PLTR$98+4.1%AI IDX4,821+1.9%
PKT
SEEDRefresh Models Deals Regulatory Sources

Chinese Open-Weight Labs Close In on Frontier Models Across Vision and Coding

DeepSeek ships a multimodal model it says rivals Opus 4.8 while Qwen 3.8 27B impresses on consumer GPUs, tightening the open-weight gap to closed…

Vector Wire — AI-assisted editorial illustration

ANALYSIS DeepSeek's launch of a multimodal model it says rivals Anthropic's Opus 4.8, paired with practitioner reports of Qwen 3.8's strong performance on consumer hardware, signals that Chinese open-weight labs are compressing the capability distance to frontier closed models on two axes — vision and agentic coding — simultaneously.

Why it matters

The multimodal frontier has been dominated by closed-model providers charging premium API prices. DeepSeek debuted V4 Flash Vision Exp, a new addition to its flagship V4 series that can understand visual prompts, saying the tool nears the performance of Anthropic's Opus 4.81,2. If that claim holds under independent evaluation, it marks a notable moment: an open-weight-leaning Chinese lab matching a top-tier Western closed model on multimodal tasks. Meanwhile, Alibaba's Qwen 3.8 27B is drawing attention from the local-inference community for agentic coding strength at aggressive quantization levels, suggesting the reasoning layer is advancing in parallel3,4.

The big picture

DeepSeek unveiled the experimental model as a paid-API-only offering on its developer platform. The company has open-sourced many of its earlier models, and may release a free version of V4 Flash Vision Exp later on. One user reported running V4 Flash on DeepInfra, describing it as "really cheaper than the official API with no peak hours"5. That combination — a paid launch with a history of subsequent open-sourcing, plus third-party hosting already undercutting the official API — illustrates how quickly open-weight economics erode pricing power even before weights are formally released.

On the reasoning side, Qwen 3.8 27B is generating practitioner enthusiasm at quantization levels previously considered too lossy for serious work. One user reported that Qwen 3.8 27B at Q6 maintained a speed of around 60–63 tokens per second throughout nearly 20 hours of non-stop goal-oriented agentic coding work across an RTX 3090 and an RTX 3060. A different user, running the model at the far more aggressive Q3_xxs quantization on an RTX 4060 Ti 16GB, reported it "one shot multiple serious coding tasks, resulting in fully working games or web apps". That user said Qwen 3.6 35B, the prior generation, "either completely failed on some of these or struggled a lot and needed hours/days of assistance/prompting, feedback to make it work" on the same tasks.

ANALYSIS The performance-per-parameter story is as important as the headline capability claims. The Q3_xxs user reported speeds of 30–35 tokens per second when fully in VRAM, dropping to 21–22 tokens per second at long context. That user said older dense models like Gemma 3 27B and Mistral small 24B only managed 13–17 tokens per second at best. That speed differential at comparable parameter counts suggests architectural efficiency gains in the latest Qwen generation, not merely scale.

The Q3_xxs user also noted limitations: the model "sometimes misunderstands things during regular convos or fails at basic sorting or counting few scores, while one shotting serious math/logic tasks". The user speculated this could stem from the low quantization level or from the model being optimized for code. That profile — strong on structured reasoning, weaker on casual conversation — is consistent with a model tuned for agentic workflows rather than general chat.

DeepSeek's multimodal move carries a different strategic implication. By benchmarking V4 Flash Vision Exp against Opus 4.8 specifically, DeepSeek is positioning itself not against other open-weight competitors but against the closed frontier. The framing is deliberate: it invites direct comparison with Anthropic's flagship and implicitly argues that the open-weight ecosystem can match proprietary performance.

The immediate question for V4 Flash Vision Exp is whether independent benchmarks confirm DeepSeek's self-reported proximity to Opus 4.8. DeepSeek's track record of open-sourcing earlier models suggests a weights release could follow, which would let the community stress-test the claim directly. For Qwen 3.8, the practitioner reports are early and anecdotal, but they point to a model that may reshape expectations for what consumer GPUs can run in agentic coding pipelines. Taken together, these two releases suggest that the next phase of the open-weight race will be fought on multimodal capability and inference efficiency simultaneously — and that Chinese labs intend to compete on both fronts at once.