GGPONC: A Corpus of German Medical Text with Rich Metadata Based on Clinical Practice Guidelines

The lack of publicly accessible text corpora is a major obstacle for progress\nin natural language processing. For medical applications, unfortunately, all\nlanguage communities other than English are low-resourced. In this work, we\npresent GGPONC (German Guideline Program in Oncology NLP Corpus), a freely\ndistributable German language corpus based on clinical practice guidelines for\noncology. This corpus is one of the largest ever built from German medical\ndocuments. Unlike clinical documents, clinical guidelines do not contain any\npatient-related information and can therefore be used without data protection\nrestrictions. Moreover, GGPONC is the first corpus for the German language\ncovering diverse conditions in a large medical subfield and provides a variety\nof metadata, such as literature references and evidence levels. By applying and\nevaluating existing medical information extraction pipelines for German text,\nwe are able to draw comparisons for the use of medical language to other\ncorpora, medical and non-medical ones.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC