CSFCube -- A Test Collection of Computer Science Research Articles for Faceted Query by Example
Query by Example is a well-known information retrieval task in which a\ndocument is chosen by the user as the search query and the goal is to retrieve\nrelevant documents from a large collection. However, a document often covers\nmultiple aspects of a topic. To address this scenario we introduce the task of\nfaceted Query by Example in which users can also specify a finer grained aspect\nin addition to the input query document. We focus on the application of this\ntask in scientific literature search. We envision models which are able to\nretrieve scientific papers analogous to a query scientific paper along\nspecifically chosen rhetorical structure elements as one solution to this\nproblem. In this work, the rhetorical structure elements, which we refer to as\nfacets, indicate objectives, methods, or results of a scientific paper. We\nintroduce and describe an expert annotated test collection to evaluate models\ntrained to perform this task. Our test collection consists of a diverse set of\n50 query documents in English, drawn from computational linguistics and machine\nlearning venues. We carefully follow the annotation guideline used by TREC for\ndepth-k pooling (k = 100 or 250) and the resulting data collection consists of\ngraded relevance scores with high annotation agreement. State of the art models\nevaluated on our dataset show a significant gap to be closed in further work.\nOur dataset may be accessed here: https://github.com/iesl/CSFCube\n
Paper
References (81)
Scroll for more · 38 remaining