Advances in object recognition flourished in part because of the availability\nof high-quality datasets and associated benchmarks. However, these\nbenchmarks---such as ILSVRC---are relatively task-specific, focusing\npredominately on predicting class labels. We introduce a publicly-available\ndataset that embodies the task-general capabilities of human perception and\nreasoning. The Human Similarity Judgments extension to ImageNet (ImageNet-HSJ)\nis composed of human similarity judgments that supplement the ILSVRC validation\nset. The new dataset supports a range of task and performance metrics,\nincluding the evaluation of unsupervised learning algorithms. We demonstrate\ntwo methods of assessment: using the similarity judgments directly and using a\npsychological embedding trained on the similarity judgments. This embedding\nspace contains an order of magnitude more points (i.e., images) than previous\nefforts based on human judgments. Scaling to the full 50,000 image set was made\npossible through a selective sampling process that used variational Bayesian\ninference and model ensembles to sample aspects of the embedding space that\nwere most uncertain. This methodological innovation not only enables scaling,\nbut should also improve the quality of solutions by focusing sampling where it\nis needed. To demonstrate the utility of ImageNet-HSJ, we used the similarity\nratings and the embedding space to evaluate how well several popular models\nconform to human similarity judgments. One finding is that more complex models\nthat perform better on task-specific benchmarks do not better conform to human\nsemantic judgments. In addition to the human similarity judgments, pre-trained\npsychological embeddings and code for inferring variational embeddings are made\npublicly available. Collectively, ImageNet-HSJ assets support the appraisal of\ninternal representations and the development of more human-like models.\n