Bio-SODA: Enabling Natural Language Question Answering over Knowledge\n Graphs without Training Data

The problem of natural language processing over structured data has become a\ngrowing research field, both within the relational database and the Semantic\nWeb community, with significant efforts involved in question answering over\nknowledge graphs (KGQA). However, many of these approaches are either\nspecifically targeted at open-domain question answering using DBpedia, or\nrequire large training datasets to translate a natural language question to\nSPARQL in order to query the knowledge graph. Hence, these approaches often\ncannot be applied directly to complex scientific datasets where no prior\ntraining data is available.\n In this paper, we focus on the challenges of natural language processing over\nknowledge graphs of scientific datasets. In particular, we introduce Bio-SODA,\na natural language processing engine that does not require training data in the\nform of question-answer pairs for generating SPARQL queries. Bio-SODA uses a\ngeneric graph-based approach for translating user questions to a ranked list of\nSPARQL candidate queries. Furthermore, Bio-SODA uses a novel ranking algorithm\nthat includes node centrality as a measure of relevance for selecting the best\nSPARQL candidate query. Our experiments with real-world datasets across several\nscientific domains, including the official bioinformatics Question Answering\nover Linked Data (QALD) challenge, show that Bio-SODA outperforms publicly\navailable KGQA systems by an F1-score of least 20% and by an even higher factor\non more complex bioinformatics datasets.\n

Paper

References (45)

Scroll for more · 33 remaining

Similar papers

© 2026 NYSGPT2525 LLC