Using publicly available data to determine the performance of methodological\ncontributions is important as it facilitates reproducibility and allows\nscrutiny of the published results. In lung nodule classification, for example,\nmany works report results on the publicly available LIDC dataset. In theory,\nthis should allow a direct comparison of the performance of proposed methods\nand assess the impact of individual contributions. When analyzing seven recent\nworks, however, we find that each employs a different data selection process,\nleading to largely varying total number of samples and ratios between benign\nand malignant cases. As each subset will have different characteristics with\nvarying difficulty for classification, a direct comparison between the proposed\nmethods is thus not always possible, nor fair. We study the particular effect\nof truthing when aggregating labels from multiple experts. We show that\nspecific choices can have severe impact on the data distribution where it may\nbe possible to achieve superior performance on one sample distribution but not\non another. While we show that we can further improve on the state-of-the-art\non one sample selection, we also find that on a more challenging sample\nselection, on the same database, the more advanced models underperform with\nrespect to very simple baseline methods, highlighting that the selected data\ndistribution may play an even more important role than the model architecture.\nThis raises concerns about the validity of claimed methodological\ncontributions. We believe the community should be aware of these pitfalls and\nmake recommendations on how these can be avoided in future work.\n