Concept Extraction to Identify Adverse Drug Reactions in Medical Forums: A Comparison of Algorithms

Social media is becoming an increasingly important source of information to\ncomplement traditional pharmacovigilance methods. In order to identify signals\nof potential adverse drug reactions, it is necessary to first identify medical\nconcepts in the social media text. Most of the existing studies use\ndictionary-based methods which are not evaluated independently from the overall\nsignal detection task.\n We compare different approaches to automatically identify and normalise\nmedical concepts in consumer reviews in medical forums. Specifically, we\nimplement several dictionary-based methods popular in the relevant literature,\nas well as a method we suggest based on a state-of-the-art machine learning\nmethod for entity recognition. MetaMap, a popular biomedical concept extraction\ntool, is used as a baseline. Our evaluations were performed in a controlled\nsetting on a common corpus which is a collection of medical forum posts\nannotated with concepts and linked to controlled vocabularies such as MedDRA\nand SNOMED CT.\n To our knowledge, our study is the first to systematically examine the effect\nof popular concept extraction methods in the area of signal detection for\nadverse reactions. We show that the choice of algorithm or controlled\nvocabulary has a significant impact on concept extraction, which will impact\nthe overall signal detection process. We also show that our proposed machine\nlearning approach significantly outperforms all the other methods in\nidentification of both adverse reactions and drugs, even when trained with a\nrelatively small set of annotated text.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC