The Bio Report
The Bio Report

Leveraging the Power of Vocal Biomarkers

Physicians can learn a lot from talking to patients, and not just from the words they say, but from millions of data points from acoustic features of their speech, such as pitch, vocal cord vibration patterns, and micro-instabilities in the voice. Canary Speech has developed an AI-based diagnostic l

Featured Speakers

Levine Media Group HostKang Shu Guest

Topics Discussed

Episode Summary

Executive Summary: The episode explores Canary Speech’s AI vocal biomarker platform, which analyzes how people speak—not what they say—to screen for neurological, psychiatric, and cognitive conditions. Chief Medical Officer Kang Shu explains how the tool detects disease-linked speech signals, integrates into ambient clinical workflows and telehealth, supports early screening and reimbursement, and aims to expand into additional diseases while remaining consent-based and clinician-friendly.

Main Topics: Voice-based disease detection (Priority: 5/5): Canary Speech uses speech physiology—pitch, cadence, vibration, jitter, and other vocal biomarkers—to detect patterns associated with depression, anxiety, Alzheimer’s, Parkinson’s, and other conditions. Clinical diagnosis challenges and value of early detection (Priority: 5/5): Shu outlines how mental health and cognitive disorders are currently diagnosed, the time burden of screening, and why earlier detection could improve outcomes and prevent deterioration. Machine learning validation and biomarker extraction (Priority: 4/5): The company trains models on patient and healthy-control data, extracts features from audio via signal processing, and validates performance using test sets and multiple ML methods. Workflow integration and product use cases (Priority: 5/5): Canary Speech is embedded invisibly into ambient scribe software, telehealth systems, hospital monitoring, and wearable/remote monitoring to minimize disruption for clinicians. Regulatory, privacy, and consent considerations (Priority: 4/5): Shu emphasizes that the product is a clinical decision support tool rather than a diagnostic, is HIPAA/GDPR compliant, and requires patient consent when used in clinical settings. Commercial model and physician adoption (Priority: 4/5): The company sells B2B SaaS, has reimbursable CPT code pathways, and positions the product as time-saving and ROI-positive to overcome physician resistance to new technology. Future indication expansion and fundraising (Priority: 3/5): Canary Speech is adding new conditions such as multiple sclerosis and exploring PTSD, ADHD, autism, pain, lung cancer, COPD, and CHF, backed by venture funding and enterprise partnerships.

Key Arguments: Speech contains rich physiological signals that can reveal disease states even when the words themselves are irrelevant. Early detection of depression, anxiety, and dementia can improve treatment timing and potentially prevent severe harm or disease progression. The system does not analyze content or meaning; it focuses on how speech is produced, making it less subjective than self-report. Robust validation across accents, geographies, ages, sexes, and languages is essential before deployment. Canary Speech is designed to fit into existing clinical workflows invisibly, reducing adoption friction for physicians. As a clinical decision support tool, it can prompt further testing rather than replace diagnosis. The business case includes both reimbursable screening codes and downstream revenue/quality benefits from earlier condition capture. The company’s technology moat is strengthened by patents and enterprise partnerships, which help attract investors.

Data Points: Speech elements analyzed per 10 milliseconds: 2,548 - Shu says the system tracks distinct speech-production elements at this cadence. Speech sample needed: About 45 seconds - Minimum amount of speech Canary Speech needs to process. Data elements per minute: About 15 million - Approximate number of features collected from a one-minute speech sample. Typical words per minute: About 110 words - Compared with the far larger number of acoustic data points gathered. Training/test split: About 80% training set, remainder test set - How the company builds and validates its models. Result latency: 10 seconds or less - Time for clinicians to see results after speech is processed. Performance range: High 70s to up to 97% - Shu cites sensitivity/specificity ranges depending on disease; Parkinson’s is mentioned at the high end. Patents: 17 issued, 12 pending - Evidence of the company’s technology moat. Languages validated: Japanese and Spanish - Languages in which the models have been validated. Enterprise partnerships: Microsoft, Zoom, Samsung - Partnerships cited as helping attract investors and support deployment.

Pivotal Quotes: "We don't analyze what you're saying, so the words don't matter to us." — Kang Shu: Explaining that the platform focuses on acoustic biomarkers rather than language content. "We actually look at the how because I can't really change the how I'm saying it." — Kang Shu: Clarifying why vocal physiology can be more informative than self-reported words. "If you can show that it actually helps to benefit patient care... and give them more time when they go home and have less pajama time, those things backed up by rigorous data, you usually can get some adoption." — Kang Shu: Describing what drives physician adoption of new clinical technology.

Implications: Voice-based AI could make screening for mental health, cognitive decline, and other diseases faster, cheaper, and more routine inside existing workflows. If validated broadly, it may shift care toward earlier, passive detection and longitudinal monitoring.

🔓 Sign Up for Unlimited Episode Search

About The Bio Report

The Bio Report podcast, hosted by award-winning journalist Daniel Levine, focuses on the intersection of biotechnology with business, science, and policy.

View all episodes from The Bio Report