Text2Cypher: Bridging Natural Language and Graph Databases

Knowledge graphs use nodes, relationships, and properties to represent\narbitrarily complex data. When stored in a graph database, the Cypher query\nlanguage enables efficient modeling and querying of knowledge graphs. However,\nusing Cypher requires specialized knowledge, which can present a challenge for\nnon-expert users. Our work Text2Cypher aims to bridge this gap by translating\nnatural language queries into Cypher query language and extending the utility\nof knowledge graphs to non-technical expert users.\n While large language models (LLMs) can be used for this purpose, they often\nstruggle to capture complex nuances, resulting in incomplete or incorrect\noutputs. Fine-tuning LLMs on domain-specific datasets has proven to be a more\npromising approach, but the limited availability of high-quality, publicly\navailable Text2Cypher datasets makes this challenging. In this work, we show\nhow we combined, cleaned and organized several publicly available datasets into\na total of 44,387 instances, enabling effective fine-tuning and evaluation.\nModels fine-tuned on this dataset showed significant performance gains, with\nimprovements in Google-BLEU and Exact Match scores over baseline models,\nhighlighting the importance of high-quality datasets and fine-tuning in\nimproving Text2Cypher performance.\n

Paper

References (37)

Scroll for more · 25 remaining

Similar papers

© 2026 NYSGPT2525 LLC