Skip to content
VECTOR WIREAI INTELLIGENCE
UTC
Home

Disaggregated Inference Moves from Design Pattern to Production Default

OpenAI, vLLM, and real-world multi-agent workloads show disaggregated inference serving is becoming the production default, bringing new efficiency and…

OpenAI's inference team is splitting routing into a control plane and data plane2. The open-source vLLM project is evolving its prefill-decode protocol with contributions from AWS and Red Hat. A multi-agent simulation is processing around 5 billion tokens a day on hosted DeepSeek Flash for under US$100 in daily model fees3. ANALYSIS Taken together, these developments show that disaggregated inference serving, breaking prefill, decode, and KV-cache management into independently scheduled stages, is becoming an operational reality, one that carries security implications the ecosystem is only beginning to address.

Why it matters

Monolithic inference engines bundle prompt processing, token generation, and memory management into a single process. Disaggregation breaks those apart so each can scale on different hardware, at different rates, under different policies. The payoff is efficiency and flexibility; the cost is a wider attack surface and new coordination overhead. The evidence emerging this month shows the industry working through both sides of that trade-off at once.

The big picture

OpenAI's inference team has moved from reactive, signal-driven load balancing to an explicit policy architecture split into a control plane and a data plane. Lou, who works on OpenAI's inference team, described the evolution: the system began by "routing based on feedback loops driven by engine signals" and shifted to "a more explicit and a predictable policy which is still informed by engine signals". That language maps directly onto disaggregation: separating the decision about where a request goes (control plane) from the mechanics of moving tokens (data plane).

The open-source vLLM project is pursuing the same structural split. Nikola, a research engineer at Mistral AI and vLLM maintainer, will present joint work with AWS and Red Hat at PyTorch North America on "how VLMs disaggregated serving has evolved to support the latest generation of hybrid models," including developments to the prefill-decode (PD) protocol for better end-to-end latency and new KV pinning mechanisms for reliability at scale. Separately, Arkadeep from PyTorch engineering at Red Hat will present on "when and why you should disaggregate your LLM serving" to maximize inference performance based on hardware and cluster configuration. ANALYSIS The convergence of a frontier lab, an open-source engine, and an enterprise Linux vendor on the same architectural pattern suggests disaggregation is becoming the default design for production serving.

Meanwhile, a concrete demand-side case illustrates why the economics matter. Slow Vale, an LLM-driven life simulation, runs 800-plus persistent AI agents on a single continuously running server. Each character makes roughly 300 to 400 LLM calls per day with an average context of around 30,000 tokens per call. The Chinese server processes around 5 billion tokens a day. The runtime uses hosted DeepSeek Flash. On October 7, 2026, actual usage was approximately 4.424 billion tokens across 163,742 requests, costing CNY 472.33. The developer noted that the Chinese server's model fees are below US$100 per day. ◆ A workload of that shape, high concurrency, long shared contexts, and bursty per-agent calls, is precisely the profile that benefits most from disaggregated prefill and KV-cache reuse, because many agents share overlapping context windows that need not be recomputed.

Between the lines

Disaggregation creates new trust boundaries, and a medium-severity vulnerability disclosed in vLLM shows what happens when those boundaries are not enforced. CVE-2026-105754 documents that vLLM's scale-out transport splits a multimodal request into a trusted render step and a separate generate step1. The generate route "decodes a caller-supplied features object" containing serialized encoder tensors, multimodal hashes, and placeholder ranges, then "forwards it into the engine as if it had come from the trusted renderer, with no rebinding to (or validation against) the active model's renderer contract". The vulnerability affects versions at or below 0.25.1 and is patched in 0.30.0. ◆ The flaw is architecturally instructive: once prefill and decode become separate HTTP endpoints, the contract between them must be cryptographically or structurally enforced, not assumed. Every team adopting disaggregated serving inherits this class of risk.

OpenAI's talk also covers "protection mechanisms that help keep the system stable" under production stress. ◆ The pairing of routing policy with explicit stability guardrails reinforces the point: disaggregation demands new safety layers that monolithic engines never needed.

What's next

PyTorch North America, scheduled for October 20 and 21 in San Jose, will feature at least two presentations on disaggregated serving. The vLLM work with AWS and Red Hat on hybrid-model support and KV pinning will detail how the PD protocol is adapting to what Nikola called "the latest generation of hybrid models". For operators running high-concurrency workloads at the scale Slow Vale demonstrates, the practical question is whether open-source tooling can match the bespoke control-plane and data-plane separation that OpenAI has built internally. The vLLM vulnerability patch in version 0.30.0 sets a baseline: disaggregation without validated inter-stage contracts is disaggregation with an open door.