FlexiDataGen: An Adaptive LLM Framework for Dynamic Semantic Dataset Generation in Sensitive Domains

Dataset availability and quality remain critical challenges in machine learning, especially in domains where data are scarce, expensive to acquire, or constrained by privacy regulations. Fields such as healthcare, biomedical research, and cybersecurity frequently encounter high data acquisition costs, limited access to annotated data, and the rarity or sensitivity of key events. These issues, collectively referred to as the dataset challenge, hinder the development of accurate, generalizable machine learning models in high-stakes domains.To address this challenge, we introduce FlexiDataGen, an adaptive large language model (LLM) framework designed for dynamic semantic prompt-dataset generation in sensitive domains. FlexiDataGen autonomously synthesizes rich, semantically coherent, and linguistically diverse prompt datasets that can be used to construct task-specific datasets for downstream model training, evaluation, and adaptation. The framework integrates four core components: (1) syntactic-semantic analysis, (2) retrieval-augmented generation, (3) dynamic element injection, and (4) iterative paraphrasing with semantic validation. Together, these components ensure the generation of high-quality, domain-relevant prompt data.Experimental results demonstrate that FlexiDataGen effectively alleviates data scarcity and annotation bottlenecks by enabling scalable and controlled generation of prompt datasets, supporting robust and privacy-preserving machine learning development in sensitive domains.

Paper

Similar papers

© 2026 NYSGPT2525 LLC