An Evaluation of Two Commercial Deep Learning-Based Information Retrieval Systems for COVID-19 Literature
The COVID-19 pandemic has resulted in a tremendous need for access to the\nlatest scientific information, primarily through the use of text mining and\nsearch tools. This has led to both corpora for biomedical articles related to\nCOVID-19 (such as the CORD-19 corpus (Wang et al., 2020)) as well as search\nengines to query such data. While most research in search engines is performed\nin the academic field of information retrieval (IR), most academic search\nengines$\\unicode{x2013}$though rigorously evaluated$\\unicode{x2013}$are\nsparsely utilized, while major commercial web search engines (e.g., Google,\nBing) dominate. This relates to COVID-19 because it can be expected that\ncommercial search engines deployed for the pandemic will gain much higher\ntraction than those produced in academic labs, and thus leads to questions\nabout the empirical performance of these search tools. This paper seeks to\nempirically evaluate two such commercial search engines for COVID-19, produced\nby Google and Amazon, in comparison to the more academic prototypes evaluated\nin the context of the TREC-COVID track (Roberts et al., 2020). We performed\nseveral steps to reduce bias in the available manual judgments in order to\nensure a fair comparison of the two systems with those submitted to TREC-COVID.\nWe find that the top-performing system from TREC-COVID on bpref metric\nperformed the best among the different systems evaluated in this study on all\nthe metrics. This has implications for developing biomedical retrieval systems\nfor future health crises as well as trust in popular health search engines.\n