Etched doubled its valuation to $21 billion2, Velaura AI crossed the $1 billion mark3,6, and Cerebras began shipping a new rack that doubles inference throughput1,5. ANALYSIS The convergence is not coincidental: as AI workloads tilt from training to serving, investors and chip designers alike are repricing the stack around tokens per second per dollar, not peak FLOPS.
The simultaneous surge of capital into three distinct inference-silicon architectures — Etched's ASIC, Velaura's low-power approach, and Cerebras's wafer-scale engine — signals that the market no longer views Nvidia's general-purpose GPUs as the only viable path for serving AI at scale.
Etched raised $700 million at a $21 billion valuation led by Jane Street, which both tested and purchased the startup's hardware7. That round came just weeks after a $300 million raise backed by Nvidia at a $10.3 billion valuation4. Separately, Velaura AI closed a $110 million Series A led by Seligman Ventures, with Samsung Catalyst Fund and Mayfield among the backers, pushing its valuation past $1 billion. And Cerebras unveiled CS-4, a server rack powered by three WSE-3 Turbo chips built around its new Nexus architecture, with first shipments starting this quarter.
ANALYSIS Each of these companies is attacking inference from a different angle, but all three share a thesis: the inference bottleneck is architectural, not just a matter of adding more GPUs.
Etched's co-founder and COO Robert Wachen framed the problem in two stages. "Inference is built in two stages," Wachen said, "prefill and decode". Etched designed separate silicon for each: a low-voltage prefill chip that packs in more transistors without typical heat problems, and a new memory-and-interconnect system it calls "cluster-scale memory" for the decode phase. The company delivers its technology as full systems it calls "frontier inference clusters".
Cerebras took a different route — extracting more from existing silicon. The CS-4 uses the same 5nm WSE-3 wafer as the CS-3 but doubles clock speeds through improved power delivery and cooling. Each rack now holds three wafers, up from two, and retains 44GB of SRAM capacity per wafer with 43 PB/s of total on-chip memory bandwidth. Off-wafer I/O doubles to 2.4Tb/s from 1.2Tb/s. Cerebras claims the result is up to 30x interactivity improvement compared with GPUs.
Velaura, the earliest-stage of the three, is targeting low-power chips for data centers and physical AI applications such as robotics. Its $110 million raise at a billion-dollar-plus valuation on a Series A underscores how aggressively investors are funding inference alternatives even at the pre-revenue stage.
ANALYSIS The velocity of Etched's valuation trajectory — from $5 billion in December to $10.3 billion in July to $21 billion in August — reflects something beyond hype: Jane Street, a quantitative trading firm with exacting performance standards, led the latest round after becoming a customer. That a buyer-turned-investor is leading the deal at this price lends the valuation a demand-side anchor that pure financial investors cannot provide.
Cerebras's CS-4 strategy reveals a complementary bet. Rather than designing new silicon, the company doubled performance at what it describes as roughly the same bill-of-materials cost per wafer. A new field-upgradeable I/O module enables heterogeneous, disaggregated inference architectures — pairing wafer-scale SRAM with external HBM-based systems to address memory capacity constraints. The modular "backpack" rack design, which separates power, cooling, I/O, and compute into independently serviceable modules, is engineered to compress deployment timelines.
Taken together, the three announcements suggest that inference silicon is fragmenting into specialized niches — high-throughput ASIC clusters, wafer-scale SRAM engines, and low-power edge-to-datacenter designs — rather than converging on a single architecture. Capital is flowing to all of them simultaneously.
What's next
Cerebras has said CS-4 shipments begin this quarter, providing the first real-world performance data against its claims. Cerebras has outlined a roadmap targeting 20x throughput improvement by 2027 and has already co-designed the next-generation wafer-scale engine alongside the Nexus platform. Etched, having now raised over $1 billion across its July and August rounds, will face pressure to convert Jane Street's endorsement into a broader customer roster. ANALYSIS For the inference chip market, the question is no longer whether alternatives to general-purpose GPUs will attract capital — it is whether the deployed performance will justify the valuations now being assigned.