Invest Like the Best with Patrick O'Shaughnessy
Invest Like the Best with Patrick O'Shaughnessy

Chetan Puttagunta and Modest Proposal - Capital, Compute & AI Scaling - [Invest Like the Best, EP.400]

My guests today are Chetan Puttagunta and Modest Proposal. Chetan is a General Partner at venture firm Benchmark, while Modest Proposal is an anonymous guest who manages a large pool of capital in the public markets. Both are good friends and frequent guests on the show, but this is the first time t

Featured Speakers

Chetan Pudagunta GuestModest Proposal Guest

Topics Discussed

Episode Summary

Executive Summary: Patrick O'Shaughnessy speaks with Benchmark's Chetan Pudagunta and anonymous investor Modest Proposal about AI's shift from pre-training to test-time compute, synthetic data limits, and how that changes model strategy, hyperscaler CapEx, and where venture/public market value may accrue.

Main Topics: Pre-training hits a plateau (Priority: 10/5): The guests argue text-data and synthetic-data limits are slowing brute-force pre-training gains. Test-time compute becomes the new scaling axis (Priority: 10/5): AI progress is shifting toward reasoning, verification, and inference-time problem solving. Model layer gets cheaper and more open (Priority: 9/5): Small teams can now reach frontier performance using open-source models and fine-tuning. Public market implications (Priority: 9/5): CapEx, cloud architecture, power demand, and valuations may need re-rating if inference dominates. Application-layer acceleration (Priority: 10/5): AI apps are unlocking clear ROI, shorter sales cycles, and new enterprise categories. Big model players and strategic moats (Priority: 8/5): OpenAI, Anthropic, xAI, Google, and Meta face different distribution and defensibility paths. AGI/ASI and recursive self-improvement (Priority: 7/5): They debate how close AGI is, what ASI would mean, and whether machines can surpass training bounds.

Key Arguments: Pre-training scaling is plateauing because human text and synthetic data are both insufficient. Test-time compute can improve reasoning, but verifier/search limits may cap returns on more compute. Open-source Llama lets tiny teams fine-tune to frontier-like performance with far less capital. AI app economics are improving fast; some inference costs are down 100x-200x and margins can hit 95%. CapEx may re-align with usage if inference, not training, drives spend; that's better for hyperscalers. OpenAI's consumer brand and distribution may matter as much as model quality if free rivals emerge. Google may still have the ingredients to win in AI, but the payoff may not resemble search's dominance. The app layer is seeing real enterprise pull, with sales cycles compressing from months to days/weeks.

Data Points: AI spend growth: 8x year-over-year - Private-market AI software spend grew from 2023 to 2024. Application investments since ChatGPT: 25 investments - Benchmark's pace of AI investing since November 2022. Infrastructure investments: 4 companies - Of Benchmark's 25 AI investments, four were infrastructure names. Fund size: $500 million - Benchmark fund size referenced by Pudagunta. Makeup of market cap tied to AI: 40% to 45% - Anonymous investor's estimate of market cap directly exposed to AI themes. Public market earnings multiple: 24 times earnings - Used to describe the broad optimism in public equities. Google multiple: 19 or 20 times - Cited as a more moderate valuation among large tech names. Inference cost decline: 100x, 200x - Estimated drop in inference costs for application developers versus two years ago. Inference pricing: $15 to $20 per million tokens - Earlier frontier-model inference cost cited for first-wave AI apps. OpenAI consumer price: $20 - Monthly consumer subscription referenced in discussion of ChatGPT distribution. Model cluster size: 100,000 H100s - Referenced as the scale of Meta's Llama 4 training cluster. Supercluster delivery: 300,000 to 400,000 chip super cluster - Expected to be delivered by end of next year or early 2026. Stargate timeline: 2028 delivery - Hypothesized OpenAI/Microsoft data-center project timeline. CapEx concern threshold: $20 billion or $50 billion - Scale at which training commitments became harder to justify. Potential future spend: $85 billion in cash CapEx - Referenced as Microsoft's 2025 AI-related spend including leases. Inference speedup: 900+ tokens per second - Cerebras inference on Llama 3.1 405B. Inference speedup multiple: 70 or 75 times faster - Cerebras compared with GPUs for inference. AI model performance: under a million dollars - Some teams reportedly matched frontier performance on specific use cases for under $1M.

Pivotal Quotes: "we're now shifting to a new paradigm called test-time compute" — Chetan Pudagunta: Explaining the industry move away from pre-training scaling. "I think you have to start in general with animal spirits." — Modest Proposal: Framing public-market valuation optimism around ChatGPT and AI. "If you're scared of our next model being released, we're going to run you over." — Sam Altman (quoted by Modest Proposal): Used to illustrate how model cadence can pressure developers and competitors.

Implications: Investors should watch whether inference-first economics, not another pre-training breakthrough, becomes the durable default—and whether open-source pressure keeps model power from concentrating.

From the Transcript

Of was synthetic data going to enable these models to continue to scale. Everybody assumed, as you saw that line, this problem was going to really come to the forefront in 2024. And here we are. We're here and we're all trying to train on synthetic data, the large model providers. And now, as it's been reported in the press and as all these AI lab leaders have gone on the record, we're now hitting limits. Because of synthetic data. The synthetic data, as generated by the LLMs themselves, are not enabling the scaling and pre-training to continue. And so we're now shifting to a new paradigm called test-time compute. And what test-time compute is, in a very basic way, is you actually ask the LLM to look at the problem, come up with a set of potential solutions to it, and pursue multiple solutions in parallel. You create this thing called a verifier. And you pass through the solution over and over again, iteratively. And the new paradigm of scaling, if you will, the x-axis is time measured in logarithmic scale, and intelligence is on the y-scale. And that's where we are today, where it seems that almost everybody is moving to a world where we're scaling on pre-training and training to scaling on what's now being called reasoning, or that is inference.

Chetan Pudagunta · at 3:57

These kinds of companies are seeing extraordinary usage by developers, and then all the way up the stack to the applications themselves. I think just the pace of innovation, pace of commercial success is driving a lot of excitement with private investors. What is also appealing of model stability is now we can finally assume, if this sticks, that all these companies are going to be fairly capital light. Because if you're not having to spend a lot on pre-training, if you're not going to have to spend a lot on infrastructure. Because most of the hyperscalers are now going to present you with really reliable APIs at these kinds of costs. It's a great time to be in the application development business, and it's a great time to be in the application development stack. Modest, what do you think on valuations? I think you have to start in general with animal spirits. If you go back to the week before ChatGPT was released, if you go to the fall of 2022. Tech had probably just suffered its most brutal bear market since the dot-com collapse. It was arguably worse for the median tech stock than even the financial crisis. You had some of the very large growth funds down 60, 70 percent. You had the hyperscalers laying off people for the first time ever. You had CapEx cuts, you had OpEx cuts. It was a very different vibe.

Modest Proposal · at 1:02:02

Jason, was why is stability in the model layer important? I think Sam Altman gave the right answer on this, which was six months ago, he was on a podcast and said, if you're scared of our next model being released, we're going to run you over. If you're looking forward to our next model coming out, then you're in a good position. Well, if the actual reality is the next model is going to be at inference time and not pre-training, you probably have less worry about them steamrolling. So, I think everything that we're talking about in this one pack is very conducive to a favorable economic reality for the entire ecosystem, which is all the attention capital being put towards inferencing. The real concern was: do we need to spend $50,000, $200 billion to build these ever more accurate models in pre-training? Where do prices most reflect extreme?

Sam Altman · at 58:32
🔓 Sign Up for Unlimited Episode Search

About Invest Like the Best with Patrick O'Shaughnessy

Conversations with the best investors and business builders in the world.

View all episodes from Invest Like the Best with Patrick O'Shaughnessy