On The Role of Prompt Construction In Enhancing Efficacy and Efficiency of LLM-Based Tabular Data Generation
LLM-based data generation for real-world tabular data can be challenged by the lack of sufficient semantic context in feature names used to describe columns. We hypothesize that enriching prompts with even minimal contextual information, such as a brief explanation of what each feature represents can improve both the quality and efficiency of data generation. To test this, we investigate three prompt construction methods: Expert-guided, LLM-guided, and Novel-Mapping, with the latter two being automated approaches. Using the GReaT framework, our experiments show that context-enriched prompts significantly enhance the quality of the generated data while improving training efficiency. Notably, the LLM-guided method performed on par with expert-guided approaches, demonstrating its effectiveness as a scalable alternative.
Paper
References (21)
Scroll for more · 9 remaining