Google Research's Biomarker Discovery Framework organizes candidate biomarker prioritization from wearable sensor data as an iterative, human-supervised research loop, combining multi-agent hypothesis generation with statistical analysis and literature-grounded reasoning.
The pipeline runs six phases under human supervision, Google describes, with an Orchestrator agent decomposing natural-language research directives into execution plans while shared memory, a structured fact sheet, and common tools preserve traceability. Scout agents map data schemas, missingness, and clinical endpoints with target labels kept out of feature construction; Literature and Hypotheses agents ground candidate features in prior evidence; Statistical and ML agents estimate associations and adjust for multiple testing.
Critic and Defender agents stress-test candidates for target leakage, overfitting, confounding sensitivity, and physiological implausibility, assigning reporting labels such as screened, conditional, exploratory, rejected, and unstable from a structured 11-check battery. In the mental health domain the framework used the DWB cohort and proposed sleep-timing variability features, estimating an association between sleep-duration variability and PHQ-8 severity (rho 0.252, p < 0.001); in the metabolic domain it derived a cardiovascular fitness index from steps divided by resting heart rate as a correlate of insulin resistance.
The paper reports 41 candidate digital biomarkers for mental health and 25 for metabolic outcomes across three cohorts totaling 9,279 participant-observations, with improved predictive performance alongside demographic features (delta R-squared 0.040 for depression, 0.021 for insulin resistance). On data-science and health benchmarks the framework performed competitively against per-benchmark baselines, and a blinded human expert evaluation gave it the highest mean scores across quality dimensions.
The work, led by MIT PhD student Yubin Kim during a Google internship and advised by Daniel McDuff and Hamid Palangi, is published on arXiv. The authors frame the findings as hypothesis-generating: effect sizes are modest, and the mechanisms read as literature-grounded hypotheses rather than clinical validation or causal evidence.




