A Corpus for Detecting High-Context Medical Conditions in Intensive Care Patient Notes Focusing on Frequently Readmitted Patients

A crucial step within secondary analysis of electronic health records (EHRs)\nis to identify the patient cohort under investigation. While EHRs contain\nmedical billing codes that aim to represent the conditions and treatments\npatients may have, much of the information is only present in the patient\nnotes. Therefore, it is critical to develop robust algorithms to infer\npatients' conditions and treatments from their written notes. In this paper, we\nintroduce a dataset for patient phenotyping, a task that is defined as the\nidentification of whether a patient has a given medical condition (also\nreferred to as clinical indication or phenotype) based on their patient note.\nNursing Progress Notes and Discharge Summaries from the Intensive Care Unit of\na large tertiary care hospital were manually annotated for the presence of\nseveral high-context phenotypes relevant to treatment and risk of\nre-hospitalization. This dataset contains 1102 Discharge Summaries and 1000\nNursing Progress Notes. Each Discharge Summary and Progress Note has been\nannotated by at least two expert human annotators (one clinical researcher and\none resident physician). Annotated phenotypes include treatment non-adherence,\nchronic pain, advanced/metastatic cancer, as well as 10 other phenotypes. This\ndataset can be utilized for academic and industrial research in medicine and\ncomputer science, particularly within the field of medical natural language\nprocessing.\n

Paper

References (23)

Scroll for more · 11 remaining

Similar papers

© 2026 NYSGPT2525 LLC