PyHard: a novel tool for generating hardness embeddings to support data-centric analysis

For building successful Machine Learning (ML) systems, it is imperative to\nhave high quality data and well tuned learning models. But how can one assess\nthe quality of a given dataset? And how can the strengths and weaknesses of a\nmodel on a dataset be revealed? Our new tool PyHard employs a methodology known\nas Instance Space Analysis (ISA) to produce a hardness embedding of a dataset\nrelating the predictive performance of multiple ML models to estimated instance\nhardness meta-features. This space is built so that observations are\ndistributed linearly regarding how hard they are to classify. The user can\nvisually interact with this embedding in multiple ways and obtain useful\ninsights about data and algorithmic performance along the individual\nobservations of the dataset. We show in a COVID prognosis dataset how this\nanalysis supported the identification of pockets of hard observations that\nchallenge ML models and are therefore worth closer inspection, and the\ndelineation of regions of strengths and weaknesses of ML models.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC