Synthesized Annotation Guidelines are Knowledge-Lite Boosters for Clinical Information Extraction
OBJECTIVE Generative information extraction using large language models (LLMs), particularly through prompting combined with few-shot learning, has become a popular method. In many ways such prompts with examples resemble the annotation guidelines long used for manual labeling of data for information extraction, and indeed studies have demonstrated the direct use of these guidelines as effective prompts. However, constructing annotation guidelines is both labor- and knowledge-intensive. Instead, this paper proposes to leverage LLMs' impressive ability to automatically create such annotation guidelines. MATERIALS AND METHODS Specifically, we propose a zero-shot hierarchical prompt engineering method that harvests the knowledge summarization and text generation capacity of LLMs to synthesize annotation guidelines to improve downstream LLMs while requiring minimal human input. RESULTS Zero-shot clinical named entity recognition benchmarks, 2012 i2b2 EVENT, 2012 i2b2 TIMEX, 2014 i2b2, and 2018 n2c2 showed improvements of 0.2% to 25.86% for Llama 3.1 and 5.82% to 16.13% for GPT-OSS in strict F1 scores from the no-guideline baseline. The LLM-synthesized guidelines showed equivalent or better performance compared to human-written guidelines by 0.23% to 10.00% in most tasks. DISCUSSION LLMs generate high-quality annotation guidelines following a consistent pattern (eg, title, entity types, examples) without human guidance, indicating that a representation of such a concept has been encoded during the pre-training. Nuances in definitions, however, still require adjustment by researchers to align with the project. CONCLUSION This study proposes a novel hierarchical prompt engineering method that requires minimal knowledge transfer from a human expert and is applicable to multiple biomedical domains.
Paper
References (24)
Scroll for more · 12 remaining