CSAW-M: An Ordinal Classification Dataset for Benchmarking Mammographic Masking of Cancer

Interval and large invasive breast cancers, which are associated with worse\nprognosis than other cancers, are usually detected at a late stage due to false\nnegative assessments of screening mammograms. The missed screening-time\ndetection is commonly caused by the tumor being obscured by its surrounding\nbreast tissues, a phenomenon called masking. To study and benchmark\nmammographic masking of cancer, in this work we introduce CSAW-M, the largest\npublic mammographic dataset, collected from over 10,000 individuals and\nannotated with potential masking. In contrast to the previous approaches which\nmeasure breast image density as a proxy, our dataset directly provides\nannotations of masking potential assessments from five specialists. We also\ntrained deep learning models on CSAW-M to estimate the masking level and showed\nthat the estimated masking is significantly more predictive of screening\nparticipants diagnosed with interval and large invasive cancers -- without\nbeing explicitly trained for these tasks -- than its breast density\ncounterparts.\n

Paper

References (35)

Scroll for more · 23 remaining

Similar papers

© 2026 NYSGPT2525 LLC