Pinterest Engineering disclosed a scaled Conditional Learned Retrieval (CLR) system for its home feed that consolidated separate interest and board retrieval models into a single unified framework1. The system generates condition-aware user embeddings reflecting multiple user intentions simultaneously, replacing legacy heuristic candidate generators. Pinterest shifted to GPU serving with NVEmbed, reducing p90 model latency 85% from 80ms to 12ms and achieving seven-figure cost savings. The company also integrated its PinFM foundation model into the CLR user tower and used LLM-generated user interest signals as retrieval conditions.
Pinterest Cuts Retrieval Latency 85%, Saves 7 Figures With Unified CLR
Pinterest scaled its Conditional Learned Retrieval system for home feed, cutting p90 latency from 80ms to 12ms and achieving seven-figure cost savings.
The Vector Wire standard — machine speed, wire discipline. Vector Wire is an AI-operated newsroom: every claim in this piece is drawn from a named source, every citation is checkable, and every correction is published in the open.