Researchers currently rely on ad hoc datasets to train automated\nvisualization tools and evaluate the effectiveness of visualization designs.\nThese exemplars often lack the characteristics of real-world datasets, and\ntheir one-off nature makes it difficult to compare different techniques. In\nthis paper, we present VizNet: a large-scale corpus of over 31 million datasets\ncompiled from open data repositories and online visualization galleries. On\naverage, these datasets comprise 17 records over 3 dimensions and across the\ncorpus, we find 51% of the dimensions record categorical data, 44%\nquantitative, and only 5% temporal. VizNet provides the necessary common\nbaseline for comparing visualization design techniques, and developing\nbenchmark models and algorithms for automating visual analysis. To demonstrate\nVizNet's utility as a platform for conducting online crowdsourced experiments\nat scale, we replicate a prior study assessing the influence of user task and\ndata distribution on visual encoding effectiveness, and extend it by\nconsidering an additional task: outlier detection. To contend with running such\nstudies at scale, we demonstrate how a metric of perceptual effectiveness can\nbe learned from experimental results, and show its predictive power across test\ndatasets.\n