VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

PyTorch Conference NA 2026 Dedicates Broad Track to vLLM Inference Stack

PyTorch Conference North America 2026, October 20–21 in San Jose, features vLLM across sessions on KV cache, disaggregated serving, hardware portability…

Vector Wire — AI-assisted editorial illustration

PyTorch Conference North America 2026, scheduled for October 20–21 in San Jose, CA, features vLLM across sessions spanning KV cache management, disaggregated serving, hardware portability, kernel optimization, Mixture-of-Experts inference, and production deployment1. Sessions cover attention backends, tiered KV cache offloading, elastic expert parallelism, and multi-stage prefix caching. Hardware-portability talks demonstrate vLLM running on Intel GPUs, Arm CPUs, IBM Spyre, AWS Trainium, and Google Cloud TPUs through backends including OpenVINO, Torch-Spyre, and TorchTPU. A Birds of a Feather discussion addresses contributing to vLLM and llm-d.

The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.