Characterizing Variation in Crowd-Sourced Data for Training Neural Language Generators to Produce Stylistically Varied Outputs

One of the biggest challenges of end-to-end language generation from meaning\nrepresentations in dialogue systems is making the outputs more natural and\nvaried. Here we take a large corpus of 50K crowd-sourced utterances in the\nrestaurant domain and develop text analysis methods that systematically\ncharacterize types of sentences in the training data. We then automatically\nlabel the training data to allow us to conduct two kinds of experiments with a\nneural generator. First, we test the effect of training the system with\ndifferent stylistic partitions and quantify the effect of smaller, but more\nstylistically controlled training data. Second, we propose a method of labeling\nthe style variants during training, and show that we can modify the style of\nthe generated utterances using our stylistic labels. We contrast and compare\nthese methods that can be used with any existing large corpus, showing how they\nvary in terms of semantic quality and stylistic control.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC