Unsupervised Learning with Jacob Effron

Ep 94: Applied Compute CEO on the Limits of RL, the New AI Hyperscaler & Why Post-Training Wins Inference

Episode Summary

Yash Patil, a former OpenAI researcher who worked on Codex and now runs Applied Compute, argues that post-training is the most overlooked part of the AI stack and the key to owning your intelligence. He explains when companies should optimize their harness and context first, and when it's worth updating the weights: the largest inference workloads, plus tasks that change over time or depend on judgment. That's the heart of why he believes post-training wins inference. The biggest workloads are both the most valuable to train and the most valuable to serve, and how a model is trained shapes how it should be served. Getting production serving running is the easier half; echoing Dylan Patel, he says GPUs plus vLLM plus OpenRouter "kind of" gets you there. Yash is candid about the limits of RL. It's a hill-climbing machine whose hardest part is defining the hill, which makes a company's private evals as worth protecting as its employees. It generalizes far less than pre-training. And continual learning remains blocked by the unsolved problem of extremely data-efficient training from sparse rewards. He lays out his bet on a new AI hyperscaler, built on GPUs the way AWS, GCP, and Azure were built on CPUs. He pictures it as an inverted pyramid with training at the bottom, where few can compete, and inference, routing, and harness stacked above. The conversation also covers whether every firm needs its own model, why Jevons paradox has made cost matter more than he expected, how models reward hack "like water," and why the US needs American open-weight models even though "blanket bans are not the answer."

Episode Notes

Yash Patil, a former OpenAI researcher who worked on Codex and now runs Applied Compute, argues that post-training is the most overlooked part of the AI stack and the key to owning your intelligence.

He explains when companies should optimize their harness and context first, and when it's worth updating the weights: the largest inference workloads, plus tasks that change over time or depend on judgment. That's the heart of why he believes post-training wins inference. The biggest workloads are both the most valuable to train and the most valuable to serve, and how a model is trained shapes how it should be served. Getting production serving running is the easier half; echoing Dylan Patel, he says GPUs plus vLLM plus OpenRouter "kind of" gets you there.

Yash is candid about the limits of RL. It's a hill-climbing machine whose hardest part is defining the hill, which makes a company's private evals as worth protecting as its employees. It generalizes far less than pre-training. And continual learning remains blocked by the unsolved problem of extremely data-efficient training from sparse rewards.

He lays out his bet on a new AI hyperscaler, built on GPUs the way AWS, GCP, and Azure were built on CPUs. He pictures it as an inverted pyramid with training at the bottom, where few can compete, and inference, routing, and harness stacked above. The conversation also covers whether every firm needs its own model, why Jevons paradox has made cost matter more than he expected, how models reward hack "like water," and why the US needs American open-weight models even though "blanket bans are not the answer."

(0:00) Intro
(1:42) The Case for Owning Your Intelligence
(4:35) OpenAI x Baseten and the Multi-Model Future
(7:19) Can Post-Training Beat Frontier Models?
(9:03) What Data Stays Out of Distribution
(11:51) Data Efficiency and Continual Learning
(14:31) When Updating Model Weights Makes Sense
(16:01) Online Training in Practice
(18:16) Lowering the Cost of Post-Training
(19:48) Does Every Company Need Its Own Model?
(24:06) RL in Non-Verifiable Domains
(26:59) How Applied Compute Picks Customers
(28:46) Building an Inference Business
(34:43) The Case for a New AI Hyperscaler
(38:21) Building RL Environments
(41:33) Quickfire