The Bio Report
The Bio Report

Building a Genetics Engine to Crack the Target Bottleneck

A chronic shortage of high‑quality targets remains one of the biggest constraints in drug discovery, even as therapeutic tools become more powerful and diverse. Regeneron is tackling that problem with its Regeneron Genetics Center, which has built a genetics‑driven discovery engine that integrates h

Featured Speakers

Levine Media Group HostAris Baras Guest

Topics Discussed

Episode Summary

Executive Summary: Aris Baras explains how Regeneron Genetic Center turns massive human genetics, proteomics, and AI into drug targets, biomarkers, and trial-enrichment tools. The center has moved from exomes to genomes, scaled through major partnerships, and found that proteomics and polygenic risk can outperform genetics alone for prediction. Human genetics-based programs appear to materially raise R&D success rates.

Main Topics: RGC’s mission: solve the target bottleneck (Priority: 5/5): The center was built to find better drug targets by using human genetics to identify protective or causal variants with clear therapeutic direction. Scaling from exomes to genomes and biobank-scale datasets (Priority: 5/5): Barris describes the evolution from an exome-first strategy driven by cost to a broader genome strategy enabled by falling sequencing costs and much larger sample sizes. Protective genetics and monogenic discovery (Priority: 5/5): Regeneron has repeatedly used rare, high-effect variants to find or validate targets, from PCSK9-like protective biology to monogenic diseases such as congenital hearing loss. Proteomics as a major new predictive layer (Priority: 4/5): Protein-level data from UK Biobank and other cohorts surprised the team by outperforming genomic data for near-term disease prediction and offering dynamic risk signals. AI/ML for variant interpretation and workflow acceleration (Priority: 4/5): AI helps sift through hundreds of millions of variants, improve functional interpretation, and speed coding, synthesis, and operational analysis across large datasets. Translating discovery into pipeline, indication expansion, and trial enrichment (Priority: 5/5): RGC data are used not just for target discovery but also for patient stratification, endpoint enrichment, and finding new indications for existing assets. Competitive advantage through scale, partnerships, and analytics (Priority: 4/5): Regeneron sees its edge in data depth, partnerships, proteomics, analytical talent, and the rapid handoff from discovery to therapeutic programs.

Key Arguments: Human genetics is the most reliable way to identify high-value targets because it reveals biology already proven in people. Scale is essential: rare variants require millions of samples before associations become convincing enough to fund drug development. Exomes were initially favored because they capture protein-coding variation and were far cheaper than genomes, but genome sequencing has become practical as costs fell. Proteomics adds dynamic, near-term disease information and can predict outcomes better than inherited DNA alone in many cases. Polygenic risk can be as clinically meaningful as monogenic disease when many small effects add up to large risk. RGC uses genetics and biomarkers to enrich clinical trials, improving event rates and reducing wasted study size. Human genetics also supports indication expansion by revealing additional diseases where a target or mechanism may work. AI is essential for interpreting variant consequences, integrating proteomics and clinical data, and accelerating internal workflows. Regeneron’s success rate is improved by starting with genetically validated targets, especially those with large effect sizes and clear direction of effect. Scale and integration are difficult to copy because they depend on data access, partnerships, talent, and a fast translation engine from discovery to development.

Data Points: Genome coverage of approved/experimental medicines: ~10% of the genome - Barris says only about 10% of the genome had been explored by approved or tried medicines when RGC began. People sequenced (current): Millions - RGC has already sequenced millions of people and is moving toward 10 million. Near-term ambition: 10 million sequenced - Barris says RGC has a confident path to 10 million people sequenced. Longer-term ambition: 50 million to 100 million - He describes a new aspiration to reach tens of millions, potentially 50M and 100M. Exome vs genome cost at launch: ~10x difference - At the start, whole-genome sequencing cost about 10 times more than exome sequencing. RGC team size: A couple hundred people - Barris credits a team of roughly 200 scientists and staff. Protective factor discoveries: >50 PCSK9-type protective factors - He says RGC has found over 50 protective genetic factors analogous to PCSK9. New pipeline programs from RGC: ~50 targets/programs - He says around 50 new targets have been put into the Regeneron pipeline. Annual additions to pipeline: ~10 new programs per year - RGC is adding about 10 new programs annually from discoveries. Proteomics pilot size: 50,000 individuals - The first major proteomics project in UK Biobank covered 50,000 people. Proteomics scale: Hundreds of thousands to millions - RGC says it moved from pilot to hundreds of thousands and is now moving to millions in proteomics. Cardiovascular outcomes trial size: 20,000 patients - Used as an example of how genetics could enrich large outcome studies. Typical placebo-arm event rate: ~10-12% - Barris notes relatively low event rates in broad cardiovascular outcome studies. High-risk subgroup estimate: ~30% heart attack risk in 3 years - Genetic/biomarker stratification can identify a much higher-risk subset. Industry clinical success rate: ~10% make it from clinic to approval - He cites the broad industry figure that 90% of clinical assets fail. Genetics-based success uplift: 2-3x higher in literature; 4-5x in RGC hands - Barris says human genetics can materially increase probability of success.

Pivotal Quotes: "What we were looking for was more of these types of examples." — Aris Baras: Explaining the founding motivation behind RGC: finding protective genetic stories like PCSK9 and CCR5. "The power that the proteomics data held was, I mean, it really amazed us." — Aris Baras: Describing the unexpected strength of proteomics for predicting disease beyond genetics alone. "It is increasing the probability of success." — Aris Baras: On how RGC’s genetics-first approach improves drug-development outcomes and R&D decision-making.

Implications: Regeneron’s model suggests future drug discovery will be more data-rich, genetics-led, and trial-enriched. Companies that can combine large cohorts, proteomics, and AI may find better targets faster and raise clinical success rates.

🔓 Sign Up for Unlimited Episode Search

About The Bio Report

The Bio Report podcast, hosted by award-winning journalist Daniel Levine, focuses on the intersection of biotechnology with business, science, and policy.

View all episodes from The Bio Report