This paper introduces PolyPersona, a generative framework that instantiates a persona-conditioned language model to synthesize realistic survey responses across multiple domains. The framework instruction-tunes compact chat models using parameter-efficient LoRA adapters and 4-bit quantization under a resource-adaptive training setup. A dialogue-formatted data pipeline that preserves persona cues to maintain consistent behavioral alignment across responses 11The implementation and evaluation code are available in an anonymized repository at https://anonymous.4open.science/r/Polypersona-1D70/. The resulting dataset comprises 3,568 responses spanning ten domains and 433 unique personas, enabling controlled instruction-tuning and systematic multi-domain evaluation. A multi-metric evaluation stack integrates BLEU, ROUGE, and BERTScore with surveyspecific metrics capturing structural, stylistic, and sentiment consistency. Results show that small models such as TinyLlama 1.1B and Phi-2 achieve performance on par with larger 7B-8B baselines (highest BLEU 0.090, ROUGE-1 0.429), highlighting the efficiency of persona-conditioned fine-tuning on compact architectures. Persona-grounded small models prove reliable and efficient for generating synthetic survey data, with open protocols ensuring bias monitoring and reproducibility.
Paper
References (56)
Scroll for more · 38 remaining