The Cognitive Revolution
The Cognitive Revolution

Approaching the AI Event Horizon? Part 2, w/ Abhi Mahajan, Helen Toner, Jeremie Harris, @8teAPi

Abhi Mahajan (@owlposting) explains how AI is reshaping biology and medicine, including foundation models to predict cancer treatment response and why he’s both skeptical and optimistic about current results. Helen Toner unpacks CSET’s “When AI Builds AI” report and why automated AI R&D is a maj

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: This live episode explored three converging fronts in AI: applying foundation models to biology and cancer care, understanding the possibility of recursive AI R&D, and assessing the geopolitical/security constraints around advanced AI. Guests argued AI bio is promising but bottlenecked by noisy, hard-to-validate biology; automated AI R&D may be strategically surprising rather than neatly predictable; and U.S.-China competition plus fragile infrastructure make governance and security far harder than many assume.

Main Topics: AI for biology and cancer: promise, limits, and Noetic’s approach (Priority: 5/5): Abhi Mahajan described Noetic AI’s foundation-model approach to predicting cancer treatment response from multimodal tumor data. He argued the most clinically valuable biology problems lack easy ground truth, making closed-loop optimization harder than in code or math, but still expects meaningful gains in trial stratification and eventually target discovery. Why biology ML often overstates its results (Priority: 5/5): The conversation emphasized that many biology papers are confounded by hidden variables, expensive validation, species/dose/time dependence, and weak causal interpretability. Mahajan argued the field often lacks the verifiable, immediate reward structure that makes AI progress in math/code so explosive. Recursive self-improvement and automated AI R&D (Priority: 5/5): Helen Toner summarized CSET’s workshop on "When AI Builds AI," focusing on how experts disagree about whether AI will fully replace human researchers or just accelerate parts of AI R&D. The main takeaways were that disagreement remains fundamental and automated AI R&D is a major source of strategic surprise. Security, infrastructure, and U.S.-China competition (Priority: 5/5): Jeremy Harris argued that the real AI risk landscape is dominated by infrastructure: chips, data centers, power grids, supply chains, and personnel security. He stressed that U.S. dependence on globally distributed, fragile hardware creates severe strategic vulnerability, especially under Taiwanese conflict or supply-chain disruption scenarios. Governance, transparency, and resilience (Priority: 4/5): The guests discussed policy measures including continuous transparency reporting, independent audits, model-release-independent oversight, and broader societal resilience measures such as cyber defense, biosurveillance, and epistemic security. The consensus was that current governance is improving but still too discretionary and too slow. Personal productivity and AI workflows (Priority: 3/5): The closing discussion covered how the hosts and guests are using AI tools for research, summarization, financial tracking, podcast clipping, and personal knowledge management. A recurring theme was that AI usefulness spikes when it clears practical hurdles and integrates into real workflows.

Key Arguments: Biology is not like code or math: even if a reward exists, it may be delayed, noisy, and hard to attribute, which limits the power of closed-loop RL-style experimentation. Many biology ML benchmarks are misleading because of hidden confounders, weak controls, and validation costs that obscure whether a result is actually useful. Foundation-model-style learning on rich human tumor data can uncover black-box biomarkers that humans may not understand but can still use clinically. AI R&D could accelerate itself, but experts disagree on whether AI will replace all human work in the loop or merely automate narrower subtasks. The biggest bottlenecks in AI progress may be infrastructure, supply chains, and security rather than just algorithms or ideas. U.S.-China competition makes purely cooperative AI governance unrealistic without very strong inspection, compute, and enforcement mechanisms. A major policy priority should be preserving optionality through secure data-center design, personnel screening, and infrastructure hardening before capabilities advance further. AI adoption often becomes real only when products cross a practical threshold; once they do, market demand can pull them into production very quickly.

Data Points: Oncology trials failure rate: 97% - Mahajan cited the failure rate to motivate AI that can better stratify patients and improve trial design. Human tumor data modalities used by Noetic: 4 - Pathology, spatial proteomics, whole-plex spatial transcriptomics, and exome sequencing. Spatial proteomics panel: 16-plex - Used to identify cell types in tumor profiling at Noetic. Spatial transcriptome scale: 19,000 genes - Noetic profiles the entire surface of a tumor to infer functional state. Clinical utility target: Phase one drugs - Mahajan said better prediction could reduce failures in early-stage drug development. AI-design drug failure reduction estimate: 5% to 10% lower failure rate - He referenced a McKinsey study suggesting AI-designed drugs may modestly improve outcomes. Test-time training cost example: $500 compute cost - Referenced the Stanford "Learning to Discover at Test Time" work that achieved new results at relatively low cost. Workshop length: 1.5 days - CSET’s closed-door workshop on automated AI R&D. Public-private gap in model capability: Unknown, but "not huge right now" - Harris said the gap between public models and internal lab models does not seem very large, though he lacks inside information. Timeline examples from hosts: 2025-2028 - The discussion repeatedly referenced near-term capability timelines for junior software developers, senior researchers, and AI R&D automation.

Pivotal Quotes: "biology has no verifiable ground truth" — Abi Mahajan: His core explanation for why AI-for-biology is much harder to validate than AI in code or math. "automated AI R&D is simply a major source of potential strategic surprise" — Helen Toner: Summarizing the CSET workshop conclusion on uncertainty around recursive self-improvement. "we got to listen to Chinese companies when they tell us that our export control policy is working" — Jeremy Harris: His argument that U.S. export controls are materially slowing Chinese AI progress.

Implications: Expect progress in AI bio, AI R&D, and AI governance to be real but uneven. The biggest near-term wins may come from better validation, secure infrastructure, and workflow integration—not from magical breakthroughs alone.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution