The TREC 2009 web ad hoc and relevance feedback tasks used a new document\ncollection, the ClueWeb09 dataset, which was crawled from the general Web in\nearly 2009. This dataset contains 1 billion web pages, a substantial fraction\nof which are spam --- pages designed to deceive search engines so as to deliver\nan unwanted payload. We examine the effect of spam on the results of the TREC\n2009 web ad hoc and relevance feedback tasks, which used the ClueWeb09 dataset.\nWe show that a simple content-based classifier with minimal training is\nefficient enough to rank the "spamminess" of every page in the dataset using a\nstandard personal computer in 48 hours, and effective enough to yield\nsignificant and substantive improvements in the fixed-cutoff precision (estP10)\nas well as rank measures (estR-Precision, StatMAP, MAP) of nearly all submitted\nruns. Moreover, using a set of "honeypot" queries the labeling of training data\nmay be reduced to an entirely automatic process. The results of classical\ninformation retrieval methods are particularly enhanced by filtering --- from\namong the worst to among the best.\n