Natural Language Programming in Medicine: Administering Evidence Based Clinical Workflows with Autonomous Agents Powered by Generative Large Language Models

Generative Large Language Models (LLMs) hold significant promise in healthcare, demonstrating capabilities such as passing medical licensing exams and providing clinical knowledge. However, their current use as information retrieval tools is limited by challenges like data staleness, resource demands, and occasional generation of incorrect information. This study assessed the potential of LLMs to function as autonomous agents in a simulated tertiary care medical center, using real-world clinical cases across multiple specialties. Both proprietary and open-source LLMs were evaluated, with Retrieval Augmented Generation (RAG) enhancing contextual relevance. Proprietary models, particularly GPT-4, generally outperformed open-source models, showing improved guideline adherence and more accurate responses with RAG. The manual evaluation by expert clinicians was crucial in validating models' outputs, underscoring the importance of human oversight in LLM operation. Further, the study emphasizes Natural Language Programming (NLP) as the appropriate paradigm for modifying model behavior, allowing for precise adjustments through tailored prompts and real-world interactions. This approach highlights the potential of LLMs to significantly enhance and supplement clinical decision-making, while also emphasizing the value of continuous expert involvement and the flexibility of NLP to ensure their reliability and effectiveness in healthcare settings.

Paper

References (14)

02Section of Health Informatics
03Task: What the model is supposed to provide a solution for - “ What is the next best step in management ”?
04Division of Genetics and Genomics, Boston Children’s Hospital, Harvard Medical School, Boston, Massachusetts, USA
05Institute of Artificial Intelligence for Digital Health, Weill Cornell Medical College, Cornell University, New York, New York, USA
06inherent LLM knowledge and agent operation. We and applicability of agentic responses through manual discuss how Natural Language Programming hasProcessing as the dominant paradigm when dealing with transmission of the written word
07Department of Environmental Medicine and Climate Science, Icahn School of Medicine at Mount Sinai, New York, New York, USA
08Department of Anesthesiology, Perioperative and Pain Medicine
09Section of Cardiovascular Medicine
10measures to prevent the model from getting stuck in place, or model that direct the model to modify its downstream output. are separated for clarity
11Department of Critical Care Medicine, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA
12Center for Outcomes Research and Evaluation, Yale-New Haven Hospital, New Haven, Connecticut, USA

Scroll for more · 2 remaining

Similar papers

© 2026 NYSGPT2525 LLC