Consistency in Large Language Models Ensures Reliable Patient Feedback Classification

Evaluating hospital service quality depends on analyzing patient satisfaction feedback. Human-led analyses of patient feedback have been inconsistent and time-consuming, while natural language processing approaches have been limited by constraints in handling large contexts. Large Language Models (LLMs) offer a potential solution, but their hallucination tendency hinders widespread adoption. Here we show that Global Consistency Assessment (GCA), a method directing LLM to produce a structured chain of thought as a logical argument and evaluate their reproducibility across two independent predictions, enhances the reliability of LLMs in patient feedback analysis without the use of fine-tuning or annotated dataset. GCA applied to GPT-4 successfully eliminated GPT-4's 16% hallucination rate, achieving a precision of 87% while keeping a recall of 75% in analyzing 100 patient feedback samples. Furthermore, this method markedly outperforms state-of-the-art models in a benchmark of 1170 feedbacks, with a precision-recall AUC of 89%, compared to the highest score of 59% with standalone models like GPT-4, Llama 3 and classical machine learning. Consistency assessment provides a reliable and scalable solution for identifying areas of improvement in hospital services and shows promise for any text classification task

Paper

Full text

PDF

Consistency in Large Language Models Ensures Reliable Patient Feedback Classification

Semantic Scholar · Computer Science · 2024

Abstract

Evaluating hospital service quality depends on analyzing patient satisfaction feedback. Human-led analyses of patient feedback have been inconsistent and time-consuming, while natural language processing approaches have been limited by constraints in handling large contexts. Large Language Models (LLMs) offer a potential solution, but their hallucination tendency hinders widespread adoption. Here we show that Global Consistency Assessment (GCA), a method directing LLM to produce a structured chain of thought as a logical argument and evaluate their reproducibility across two independent predictions, enhances the reliability of LLMs in patient feedback analysis without the use of fine-tuning or annotated dataset. GCA applied to GPT-4 successfully eliminated GPT-4's 16% hallucination rate, achieving a precision of 87% while keeping a recall of 75% in analyzing 100 patient feedback samples. Furthermore, this method markedly outperforms state-of-the-art models in a benchmark of 1170 feedbacks, with a precision-recall AUC of 89%, compared to the highest score of 59% with standalone models like GPT-4, Llama 3 and classical machine learning. Consistency assessment provides a reliable and scalable solution for identifying areas of improvement in hospital services and shows promise for any text classification task

Similar papers

© 2026 NYSGPT2525 LLC