Deep and diverse population synthesis for multi-person households using generative models with conditional inputs

Traditional methods of population synthesis produce stable and interpretable populations but cannot capture the interrelationships between household- and individual-level attributes. Recent deep learning methods offer this flexibility, yet can overfit high-dimensional attribute relationships without structural guidance and deviate from known structures. We develop a household level synthetic population generation framework that adapts the existing conditional input directed acyclic tabular generative adversarial network, or ciDATGAN, to multi person households. The framework combines household size specific data construction, directed acyclic graphs (DAG) informed dependency regularization, and conditional population inputs as deterministic anchoring to preserve intrahousehold associations. We apply the model to generate an open access synthetic population for New York State. The synthetic population includes nearly 20 million individuals and 7.5 million households in 2021. Validation against withheld benchmarks shows close agreement: joint distribution matching of Public Use Microdata Areas (PUMA), age, race, and disability against the full public use microdata sample (PUMS) yields an R-squared of 0.602 and a Jensen-Shannon distance of 0.280, pairwise Cramer's V differences between generated and benchmark records average below 0.007 at the person level, and classifier two-sample tests on complete household records yield AUC values of 0.527-0.549, close to chance level. The generated households reproduce cross member associations while increasing diversity by 10-17% over the PUMS sample and 13.2% over PopGen alone.

Paper

References (23)

Scroll for more · 11 remaining

Similar papers

© 2026 NYSGPT2525 LLC