Efficient Deployment of Conversational Natural Language Interfaces over Databases

Many users communicate with chatbots and AI assistants in order to help them\nwith various tasks. A key component of the assistant is the ability to\nunderstand and answer a user's natural language questions for\nquestion-answering (QA). Because data can be usually stored in a structured\nmanner, an essential step involves turning a natural language question into its\ncorresponding query language. However, in order to train most natural\nlanguage-to-query-language state-of-the-art models, a large amount of training\ndata is needed first. In most domains, this data is not available and\ncollecting such datasets for various domains can be tedious and time-consuming.\nIn this work, we propose a novel method for accelerating the training dataset\ncollection for developing the natural language-to-query-language machine\nlearning models. Our system allows one to generate conversational multi-term\ndata, where multiple turns define a dialogue session, enabling one to better\nutilize chatbot interfaces. We train two current state-of-the-art NL-to-QL\nmodels, on both an SQL and SPARQL-based datasets in order to showcase the\nadaptability and efficacy of our created data.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC