VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Pinterest Cuts Retrieval Latency 85%, Saves 7 Figures With Unified CLR

Pinterest scaled its Conditional Learned Retrieval system for home feed, cutting p90 latency from 80ms to 12ms and achieving seven-figure cost savings.

Vector Wire — AI-assisted editorial illustration

Pinterest Engineering disclosed a scaled Conditional Learned Retrieval (CLR) system for its home feed that consolidated separate interest and board retrieval models into a single unified framework1. The system generates condition-aware user embeddings reflecting multiple user intentions simultaneously, replacing legacy heuristic candidate generators. Pinterest shifted to GPU serving with NVEmbed, reducing p90 model latency 85% from 80ms to 12ms and achieving seven-figure cost savings. The company also integrated its PinFM foundation model into the CLR user tower and used LLM-generated user interest signals as retrieval conditions.

The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.