Science & evidence
Built on the largest voice
biomarker dataset in the world.
For most readers: the numbers below are what matter. For clinicians: every claim on this page traces to a peer-reviewed publication or an IRB-registered study, cited in full, linked to source, never paraphrased into marketing copy.
- 1.8M+ voice recordings
- Peer-reviewed research
- Class IIa SaMD pathway in progress
Pooled Bridge2AI-Voice + partner cohorts, 2019–2025
Meta-analytic base across respiratory & voice pathologies
Biomarker feature reliability across independent partitions
Dataset scale
Seven years of pooled voice biomarker data.
The Rezpio model is trained on 1.8M recordings pooled from Bridge2AI-Voice, academic partner cohorts, and public respiratory-voice corpora. Continuous growth since 2019.
Methodology
How the evidence was built, and what the numbers actually mean.
An evidence summary is only useful if a reader can trace each headline statistic back to a defensible methodological choice. Below: the four pillars behind every number on this page, explained for a mixed clinician and patient audience.
1,800,637 recordings, pooled from 35 studies (2003–2026).
The training corpus was assembled meta-analytically, not scraped. Each of the 35 contributing studies was harmonized to a common recording protocol (sampling rate, channel geometry, prompt taxonomy) and a common consent framework before pooling. Sources span Bridge2AI-Voice (NIH), academic partner cohorts, the University of Florida Perceptual Voice Database, the Saarbrücken Voice Database, and 31 other IRB-registered studies spanning 23 years of respiratory and phonatory research.
No single site contributes more than 18% of the pooled data, which limits site-specific overfitting and gives the reference distribution genuine geographic and demographic breadth.
Reference distributions established and published in the Journal of Voice.
Before any classification work, the corpus was used to publish normative acoustic and phonatory reference ranges, the healthy distributions of jitter, shimmer, HNR, cepstral peak prominence, formant dispersion, and 40+ additional features, stratified by age, sex, and smoking status. The reference paper appeared in the Journal of Voice (2022) and is what every downstream Rezpio biomarker is compared against.
This matters because it makes Rezpio's "your voice is deviating from your baseline" claim measurable against published norms, not against a private, opaque comparator.
Split-half reproducibility r = 0.9997, in plain terms.
Split-half reproducibility is a test where the recording data is randomly split into two halves, each half is analyzed independently, and the two resulting feature sets are compared. A perfectly reliable measurement returns r = 1.0. Rezpio's features return r = 0.9997 across independent partitions of 12,400 recordings.
In plain language: the measurement is extremely consistent when the same data is split and compared against itself. That is a strong signal of measurement reliability, it means the numbers Rezpio reports on a Tuesday would look almost identical if the pipeline were re-run on a different partition on Wednesday. It is a floor beneath every clinical claim on this page.
Mean AUC 0.899 across 46 biomarkers (Frontiers in Digital Health).
The Frontiers in Digital Health preprint reports a mean area under the ROC curve (AUC) of 0.899 across 46 voice biomarkers, evaluated for discriminating stable from decompensating respiratory states.
In plain language: an AUC of 0.5 is chance, an AUC of 1.0 is perfect, and 0.899 sits firmly in the range clinicians recognize as strong discriminative performance, comparable to established triage instruments. It is not a diagnostic verdict; it is evidence that the signal is real, separable, and worth acting on.
Peer-reviewed record
Publications & citations.
We link to the source. If you want the methods, the sample sizes, or the confidence intervals, read the paper, not our summary of it.
Normative acoustic and phonatory reference ranges derived from a pooled multi-cohort voice corpus
Ravi S, Okafor M, Chen L, Patel R, Voicebiota Reference Working Group.
Acoustic biomarkers of respiratory exacerbation detectable in continuous natural speech: a multi-site cohort study
Ravi S, Chen L, Okafor M, Patel R, et al.
Discriminative performance of 46 voice biomarkers for respiratory decompensation: a pooled evaluation across 35 cohorts
Chen L, Nguyen T, Weiss K, et al.
Split-half reproducibility of acoustic feature extraction across independent recording partitions (n = 12,400)
Okafor M, Ravi S, Chen L, Kaur J, et al.
Daily home voice recordings precede symptom escalation in asthma and COPD: the TACTICAS study
TACTICAS Investigators (independent consortium).
For clinicians
The methods, at a glance.
- Study design
- Multi-site prospective cohort + retrospective pooled analysis
- Reference standard
- Spirometry (FEV₁/FVC), physician-adjudicated exacerbation events
- Model class
- On-device transformer, INT8 quantized, ~9 MB footprint
- Reproducibility
- Split-half r = 0.9997 across independent partitions (n = 12,400)
- Regulatory posture
- FDA De Novo submission in preparation (Class II, respiratory monitor)
- Data governance
- HIPAA-aligned; audio never leaves device; IRB-registered cohorts
Reading this as a pulmonologist?
Start a health-system pilot with our clinical team, or reach out to explore how Rezpio fits into your practice.
