Literature-based discovery (LBD) is a methodology for generating research hypotheses by identifying hidden connections within the scientific literature. While its application has been predominantly in the field of biomedicine, particularly through the use of Medline, the largest freely available biomedical bibliographic database, traditional LBD methodologies have relied heavily on rule-based approaches. These include utilizing co-occurrences among biomedical concepts or extracting semantic relations through natural language processing (NLP) methods. This study ventures into novel territory by exploring the use of advanced tools like ChatGPT for LBD, with a focus on leveraging prompt engineering to enhance hypothesis generation. We employed Large Language Models (LLMs) such as GPT-3.5 and GPT-4 to simulate the discovery of relationships between medical concepts. The study specifically examines the effectiveness of these models in autonomously replicating well-established medical correlations and generating potentially novel hypotheses. Our preliminary findings suggest that while LLMs show promise in generating hypotheses that occasionally deviate from established medical knowledge, challenges persist in consistently directing these models to produce truly innovative and less-explored connections. The study highlights the potential of LLMs in enriching the LBD process, yet also underscores the need for cautious evaluation and further research to optimize their application in this domain.
Paper
Full text
Implementing Literature-based Discovery (LBD) with ChatGPT
Semantic Scholar · Computer Science · 2024
Abstract
Literature-based discovery (LBD) is a methodology for generating research hypotheses by identifying hidden connections within the scientific literature. While its application has been predominantly in the field of biomedicine, particularly through the use of Medline, the largest freely available biomedical bibliographic database, traditional LBD methodologies have relied heavily on rule-based approaches. These include utilizing co-occurrences among biomedical concepts or extracting semantic relations through natural language processing (NLP) methods. This study ventures into novel territory by exploring the use of advanced tools like ChatGPT for LBD, with a focus on leveraging prompt engineering to enhance hypothesis generation. We employed Large Language Models (LLMs) such as GPT-3.5 and GPT-4 to simulate the discovery of relationships between medical concepts. The study specifically examines the effectiveness of these models in autonomously replicating well-established medical correlations and generating potentially novel hypotheses. Our preliminary findings suggest that while LLMs show promise in generating hypotheses that occasionally deviate from established medical knowledge, challenges persist in consistently directing these models to produce truly innovative and less-explored connections. The study highlights the potential of LLMs in enriching the LBD process, yet also underscores the need for cautious evaluation and further research to optimize their application in this domain.