The performance gap between the best open-weights AI models from Chinese companies and closed frontier models from US tech firms has narrowed to 4.4 months, according to Mozilla's latest State of Open Source AI report, published on September 15 and shared with Ars prior to publication1.
The finding centers on Moonshot AI's Kimi K3, which scores just three points behind Anthropic's Claude Fable 5 on the Artificial Analysis Intelligence Index composite benchmark, while costing 30 percent of the closed model's price. The report argues that most organizations should default to open models for the majority of their workloads.
Where closed models still earn their premium
Mozilla CTO Raffi Krikorian told Ars that paying for a closed model is justified only for specific task profiles. "[A Closed model] earns its premium in a few places: expert professional work, high-intensity retrieval, and long context," Krikorian said. He framed the buy-versus-open decision as "workload-specific rather than organization-specific".
ANALYSIS The framing redraws the procurement calculus: rather than choosing a single vendor tier for an entire organization, teams would route tasks to open or closed endpoints based on complexity, retrieval depth, and context-window demands.
Cost arithmetic
At 30 percent of Claude Fable 5's price, Kimi K3 delivers near-parity performance on the composite index. ◆ For high-volume, routine inference workloads, that cost differential compounds quickly, giving enterprises a financial incentive to shift default traffic to open-weights alternatives and reserve closed-model spend for the narrow band of tasks Krikorian identified.
The 4.4-month lag quantifies how quickly open-weights releases are converging on the frontier. ◆ A gap measured in months rather than years compresses the window during which any closed model commands a defensible quality advantage, pressuring frontier labs to justify their pricing through differentiated capabilities rather than across-the-board superiority.
The Mozilla report lands as frontier labs are engaged in parallel coordination on safety. OpenAI, Anthropic, and Google DeepMind confirmed weeks-long AI safety talks on September 15[1]. ◆ The convergence of a shrinking capability gap and active cross-lab safety discussions creates a backdrop in which the competitive moat for closed models rests increasingly on trust, safety infrastructure, and specialized performance rather than raw benchmark leads.
Mozilla's September 15 publication date places the data as a current snapshot of the open-versus-closed landscape.