Towards Efficient Patient Recruitment for Clinical Trials: Application of a Prompt-Based Learning Model

Objectives All clinical trials face a significant bottleneck in identifying eligible participants, particularly due to the complexity of unstructured medical texts. Recent advances in natural language processing, especially the advent of transformer-based models, have shown promise in this domain. In this study, we evaluated the performance of a prompt-based large language model (LLM) for cohort selection from unstructured medical notes. Methods Medical records were annotated with Med-CAT using the Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT) ontology. For each trial eligibility criterion, we extracted sentences containing relevant annotated concepts through an ontology-driven summarization process. These summaries were then input into a prompt-based LLM (GPT-3.5-turbo), tasked with classifying eligibility criteria in a zero-shot setting. Model performance was assessed using the 2018 National Natural Language Processing Clinical Challenges (n2c2) dataset, which required the classification of 288 patients’ medical records according to 13 eligibility criteria. Results The proposed prompt-based model achieved overall micro and macro F-measures of 0.9061 and 0.8060, respectively— among the highest scores reported for this dataset. Conclusions Our results demonstrate that integrating ontology-based extractive summarization with prompt-based LLMs can substantially improve eligibility classification. The summarization step enhanced model focus and interpretability, particularly for long or ambiguous narratives. This pipeline offers a scalable and adaptable framework for clinical trial automation and has the potential for real-world integration with electronic medical record matching systems.

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC