Everything going on in AI - updated daily from 500+ sources
Synthetic Longitudinal Tabular Data Generation via Copula
Synthetic data generation is increasingly used to enable data sharing and secondary analysis while protecting participant privacy, particularly for longitudinal tabular health data, where repeated measures per subject create within-subject dependence that most synthetic data methods are not designed to preserve. Existing generative methods, particularly generative adversarial network (GAN)-based approaches, can model complex distributions, but their estimated dependence structures are often difficult to interpret and their performance may be unstable or prone to overfitting in modestly sized datasets. Here we show that eCDF-copula, a statistically rooted approach using the empirical cumulative distribution function (eCDF) and copula modeling, preserves within- and between-visit dependence structure. To handle pervasive missing data, we propose a two-stage strategy combining multiple imputation with copula-based synthesis, enabling a variance decomposition that quantifies replication variability across methods. We benchmarked the proposed approach against four established methods on two longitudinal clinical datasets spanning markedly different sample sizes (n=120 vs. n=3,612). eCDF-copula achieved resemblance and utility exceeding those of state-of-the-art synthetic data methods, while maintaining comparable privacy.
Read Original Article →