Everything going on in AI - updated daily from 500+ sources
Rigorous Female Breast Cancer Phenotyping Using the All of Us Research Program
Objectives Accurate phenotyping of cases and controls is essential for studying biological and environmental contributors to disease in large biobanks. We aimed to develop a flexible, customizable, and reproducible electronic health record (EHR)-based phenotyping framework for identifying disease cases and generating matched control cohorts for downstream analyses. Here, we developed the Phenotyping Algorithm for Cases and matched Controls using EHR-based Rules (PACER). Materials and Methods Applying PACER to the All of Us Research Program Curated Data Repository v8.0, we identified female breast cancer (BC) cases identified among participants recorded as female at birth using at least two BC-associated diagnostic Observational Medical Outcomes Partnership concept IDs documented at least 30 days apart. A one-to-one matched control cohort was generated by jointly matching on sex, age, genetic ancestry, and state-level residency. Clinical, socioeconomic, and genomic data were integrated for analysis. Results We identified 10,225 BC cases and generated a control cohort of the same size matched for key demographic characteristics. Comparison with a phecodeX-based BC cohort showed 91.03% agreement. Among cases responding to relevant survey items, 80.86% self-reported a personal history of BC, compared to 1.89% of controls. We detected an enrichment of BC-associated GWAS catalog variants, pathogenic mutations in known risk genes, and higher polygenic risk scores in cases compared to controls. Discussion and Conclusion Concordance across a phecodeX-based cohort, self-reported survey responses, and genomic analyses supports the validity of PACER-defined cohorts. PACER is publicly available and readily adaptable to other diseases, supporting future research in risk modeling and precision medicine.
Read Original Article →