MLPerf Inference v6.1 Adds First Agentic Workload, Trillion-Parameter Model
MLPerf Inference v6.1 introduces the first agentic workload benchmarked on datacenter hardware and the first model exceeding one trillion parameters, according to Lambda . The release also recorded 8.85% higher throughput on identical hardw…
Ollama Ships v0.34.2 Release Candidate With llama.cpp Updates
Ollama published v0.34.2-rc1 on GitHub, incorporating updated llama.cpp internals . The release candidate's changelog covers changes from v0.34.1 through v0.34.2-rc0 . No additional feature details or bug fixes were specified in the release…
Meta Open-Sources MXFP8 FlashAttention-4 for Blackwell, Hits 2.85 PF/s
Meta has open-sourced an MXFP8 extension of FlashAttention-4 that delivers end-to-end block-scaled attention on Nvidia Corp. Blackwell GPUs, reaching 2.85 PF/s on forward passes and 2 PF/s on backward passes for LLM shapes . On Meta's inter…
vLLM consolidates as default LLM serving layer across chips, clouds, and consumer GPUs
In a single week, vLLM surfaced as the serving runtime in contexts ranging from a single RTX 5090 running a 27-billion-parameter model to an eight-GPU Blackwell Ultra cluster hosting a 2.4-trillion-parameter open-weights model, and as the…
ALL BRIEFS
RANKED · 5 THIS DAYGoogle opens Google Home to third-party AI agents via MCP server
Google on September 16, 2026 opened early access to Home MCP, a Model Context Protocol server that lets third-party AI agents monitor, control, and query devices across the Google Home ecosystem . The rollout is limited to Google Home Premi…