Quootstrap: Scalable Unsupervised Extraction of Quotation-Speaker Pairs from Large News Corpora via Bootstrapping

We propose Quootstrap, a method for extracting quotations, as well as the\nnames of the speakers who uttered them, from large news corpora. Whereas prior\nwork has addressed this problem primarily with supervised machine learning, our\napproach follows a fully unsupervised bootstrapping paradigm. It leverages the\nredundancy present in large news corpora, more precisely, the fact that the\nsame quotation often appears across multiple news articles in slightly\ndifferent contexts. Starting from a few seed patterns, such as ["Q", said S.],\nour method extracts a set of quotation-speaker pairs (Q, S), which are in turn\nused for discovering new patterns expressing the same quotations; the process\nis then repeated with the larger pattern set. Our algorithm is highly scalable,\nwhich we demonstrate by running it on the large ICWSM 2011 Spinn3r corpus.\nValidating our results against a crowdsourced ground truth, we obtain 90%\nprecision at 40% recall using a single seed pattern, with significantly higher\nrecall values for more frequently reported (and thus likely more interesting)\nquotations. Finally, we showcase the usefulness of our algorithm's output for\ncomputational social science by analyzing the sentiment expressed in our\nextracted quotations.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC