Y Combinator Startup Podcast
Y Combinator Startup Podcast

The Fastest Path To Super Intelligence

Poetiq is a new startup founded by former DeepMind researchers that recently achieved a major jump on the ARC-AGI and Humanity's Last Exam benchmark by layering a recursive self-improvement system on top of existing models. In this episode of Lightcone, Poetiq's Founder & CEO Ian Fisch

Featured Speakers

Y Combinator HostIan Fisher Guest

Topics Discussed

Episode Summary

Executive Summary: Ian Fisher, co-founder/co-CEO of Poetic, argues that AI startups should stop fine-tuning frontier models and instead build recursive self-improving harnesses on top of them. He describes Poetic as a system that automatically generates better reasoning systems for hard problems, delivering benchmark-leading results on ArcAGI v2 and Humanities Last Exam at far lower cost and with model-upgrade compatibility.

Main Topics: Poetic’s recursive self-improvement approach (Priority: 5/5): Poetic builds a meta-system that improves the reasoning systems layered on top of LLMs, aiming to outperform raw models without retraining from scratch. Why harnesses beat fine-tuning (Priority: 5/5): The discussion contrasts expensive, brittle fine-tuning workflows with reusable agentic harnesses that survive frontier model updates and continuously improve with them. Benchmark results and proof of capability (Priority: 5/5): The team highlights leaderboard wins on ArcAGI v2 and Humanities Last Exam as evidence that their approach can solve extremely hard reasoning and knowledge tasks. Context engineering and prompt optimization limits (Priority: 4/5): The hosts and guest discuss how Poetic automates what humans often do manually in prompt/context engineering, but argues real gains come from code-based reasoning strategies, not just better prompts. Scientific and automated approach to agent design (Priority: 4/5): Poetic frames its system as an automated optimization process that can tune prompts, reasoning strategies, and agent components depending on the problem and model. Ian Fisher’s career path and AI advice (Priority: 3/5): Fisher recounts moving from a YC startup acquisition at Google into AI research, then offers advice to experiment daily with AI and push its boundaries.

Key Arguments: Recursive self-improvement is a major leverage point for AI, and Poetic claims to do it faster and cheaper than approaches that require training new models from scratch. Fine-tuning is strategically fragile for startups because the next frontier model release can obsolete a costly investment. Poetic’s harnesses are designed to sit above any model, so improvements persist across model upgrades instead of being reset each time. Automating the generation of reasoning systems can outperform manually built agents while using far less compute and labor. The biggest gains often come from code-based reasoning strategies and system design, not merely from optimizing prompts or context. Benchmarks like ArcAGI v2 and Humanities Last Exam demonstrate both reasoning improvement and deep knowledge extraction. AI builders should experiment every day and use AI to extend what they can build, rather than limiting themselves to current skill boundaries.

Data Points: ArcAGI v2 score: 54% - Poetic’s result on ArcAGI v2 official verification, surpassing Gemini 3 Deep Think’s 45%. Gemini 3 Deep Think score: 45% - Referenced as the prior top benchmark result before Poetic’s ArcAGI v2 run. ArcAGI v2 cost per problem: $32 per problem - Poetic’s reported cost for its ArcAGI v2 result. Comparative cost per problem: about $70+ per problem - Approximate cost cited for Gemini 3 Deep Think versus Poetic’s lower-cost run. Humanities Last Exam score: 55% - Poetic’s result on the expert-written benchmark, ahead of the prior state of the art. Prior Humanities Last Exam SOTA: 53.1% - Anthropic’s Claude Opus 4.6 result mentioned as the previous best. Optimization cost: less than $100K - Cost of Poetic’s Humanities Last Exam optimization run. Humanities Last Exam size: 2,500 questions - Description of the benchmark as a large set of very difficult expert-written questions. Team size: 7 people - Poetic is described as a seven-person team of research scientists and research engineers. Manual-to-strategy improvement example: 5% to 95% - Example from prior DeepMind work showing the impact of adding reasoning strategies beyond prompt optimization.

Pivotal Quotes: "The core insight that we had is that we could do recursive self-improvement far faster and cheaper than all of the other ways that people had been proposing to do this." — Ian Fisher: Explaining Poetic’s foundational thesis and why it differs from traditional model-training approaches. "What we've built is a system that can automatically generate systems for your particular problem that will always outperform the underlying language models." — Ian Fisher: Describing Poetic as a harness/meta-system that outperforms raw LLMs and stays compatible with future model upgrades. "Don't limit yourself. Like anything that you imagine, you should just try to use AI and see how far you can get with it." — Ian Fisher: Closing advice to engineers and founders on how to approach the current AI moment.

Implications: For startups, the message is to build on top of frontier models with adaptive systems, not chase expensive retraining cycles. If Poetic’s claims hold, agentic harnesses could become a durable moat in AI application development.

🔓 Sign Up for Unlimited Episode Search

About Y Combinator Startup Podcast

We help founders make something people want. The Y Combinator Podcast is where builders talk about building. From the earliest days of an idea to scaling a company that changes the world, YC partners and founders share real stories, lessons, and tactics from the frontlines.

View all episodes from Y Combinator Startup Podcast