Effective Transfer Learning for Identifying Similar Questions: Matching User Questions to COVID-19 FAQs

People increasingly search online for answers to their medical questions but\nthe rate at which medical questions are asked online significantly exceeds the\ncapacity of qualified people to answer them. This leaves many questions\nunanswered or inadequately answered. Many of these questions are not unique,\nand reliable identification of similar questions would enable more efficient\nand effective question answering schema. COVID-19 has only exacerbated this\nproblem. Almost every government agency and healthcare organization has tried\nto meet the informational need of users by building online FAQs, but there is\nno way for people to ask their question and know if it is answered on one of\nthese pages. While many research efforts have focused on the problem of general\nquestion similarity, these approaches do not generalize well to domains that\nrequire expert knowledge to determine semantic similarity, such as the medical\ndomain. In this paper, we show how a double fine-tuning approach of pretraining\na neural network on medical question-answer pairs followed by fine-tuning on\nmedical question-question pairs is a particularly useful intermediate task for\nthe ultimate goal of determining medical question similarity. While other\npretraining tasks yield an accuracy below 78.7% on this task, our model\nachieves an accuracy of 82.6% with the same number of training examples, an\naccuracy of 80.0% with a much smaller training set, and an accuracy of 84.5%\nwhen the full corpus of medical question-answer data is used. We also describe\na currently live system that uses the trained model to match user questions to\nCOVID-related FAQs.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC