The emerging pattern across AI infrastructure is unmistakable: teams are treating their own production logs, agent failures, and inference traces as the primary training signal, building closed loops that compress deployment experience into model weights on a daily cadence.
Why it matters
Fine-tuning on curated, static datasets defined the first era of model customization. What Shopify, LangChain, and open-source tooling like phantom-kv now demonstrate is a second phase, where the artifacts of live deployment (corrections, rejected outputs, successful agent runs, even refusal patterns) feed back into the model itself. ANALYSIS The economic and quality incentives are converging: Shopify estimates that serving its GraphQL agent on a frontier model would cost roughly $27 million per year, while its fine-tuned replacement costs closer to $1 million, a 96% reduction1. That gap makes the engineering investment in continual learning loops self-funding at scale.
The big picture
Shopify's case study, published on the official PyTorch blog, lays out the logic plainly. "Frontier models are general-purpose, not tailored to your product. More importantly, they do not learn from production on their own," the post states. "A user correction, rejected output, or recurring failure does not make the next response better". Shopify's answer is what it calls a "flywheel": a continual learning loop that "compresses production experience into the continuous space of the model's weights". The company's GraphQL agent handles up to 2,000 requests per minute in production, and the loop retrains daily on failures captured from that traffic.
LangChain is operationalizing the same idea for its broader user base. LangSmith Fine-Tuning, announced at Interrupt in New York, uses an open-source CLI called smithtune that "reads your LangSmith traces and turns the successful runs into a supervised fine-tuning dataset"2. Training runs on Baseten Loops with dedicated GPUs in the user's own workspace; once LangSmith evaluates the new checkpoint, smithtune deploy places it on a Baseten Dedicated Inference deployment. The pipeline is designed to be handed to a coding agent, not just a human operator.
ANALYSIS Both systems share a core architectural conviction: production knowledge that accumulates in "prompt edits, retrieval examples, routing rules, and harness code" is a symptom of frozen weights, not a solution. The fix is to move that knowledge into the weights themselves.
Between the lines
Shopify's numbers on inference efficiency reinforce why this approach compounds. Gisting compressed the agent's long system prompt from roughly 6,000 tokens to about 1,500 learned gist tokens. In load tests at 350 requests per minute, end-to-end latency dropped about 38% and time-to-first-token fell about 19%. The company reports roughly 14% fewer GPUs needed for the same traffic and about 16% more requests per second on identical hardware. Shopify serves its models through vLLM, which provides continuous batching to maintain throughput.
ANALYSIS LangChain's approach differs in one important respect: it packages the loop as infrastructure for external developers rather than an internal system. By open-sourcing smithtune and routing training through Baseten, LangChain turns every LangSmith user's trace store into a potential fine-tuning dataset. The economic model is also distinct: users pay through their own Baseten accounts for both training and inference.
The phantom-kv project, meanwhile, pushes the boundary of what counts as "learning from production" in a different direction. Rather than retraining weights, it injects a learned bank of key/value tensors (roughly 18 megabytes) into the KV cache, influencing the model "only through the input channel attention already consumes"3,4. The cache is trained offline against the model's own objective and can be loaded or unloaded per request, leaving the base model "byte-for-byte unchanged". ◆ This positions production-derived behavioral modification as a hot-swappable layer rather than a permanent checkpoint edit, a design that trades the depth of weight-level learning for reversibility and modularity.
What's next
LangSmith Fine-Tuning is in public beta, with Baseten Loops still in early access; users may need to request enablement for their workspace. Shopify's flywheel retrains daily and already serves production traffic. ◆ The competitive question is whether teams that close this loop fastest, converting deployment failures into weight updates on a daily or sub-daily cadence, build a durable quality and cost advantage over those still relying on frozen frontier models wrapped in prompt engineering. Shopify's own framing is direct: "each failure is hard-won knowledge about your product, and continual learning begins by capturing that knowledge and feeding it back into the system".