Summary
This article proposes a MIDI dataset obtained from automatic transcription and scrapping. The data can be used for training generative models and for musicological analysis. It uses an LLM-powered scrapping method and an automatic music transcription algorithm to find MIDI files for a diversity of piano pieces.
Strengths
The dataset is in fact extensive, and the data can be used to train generative music models. Together with the metadata, this could lead to a diversity of works related to music generation.
Originality: the idea of a dataset is not novel; however, a dataset this large is novel.
Quality: the paper is well written and addresses all issues related to the problem
Clarity: The only issue I found is that the fact that the dataset contains solely piano pieces should be mentioned in the title.
Significance: Highly significant, yet for a niche.
Weaknesses
The dataset is a great contribution, but it might be controversial to say that the dataset falls under "fair use". This is, first, because "fair use" is different in each country; also, nothing in the CC-BY-NC licence mentions the use of models derived from this data to generate new data, that is, although the files cannot be sold themselves, the models able to generate similar excerpts could be sold and used for commercial purposes. In this sense, a CC-BY-NC-SA would be more adequate and more coherent with the "fair use" idea, or, alternatively, terms of licencing that prohibit usage for commercial models or models that might be used in commercial distributions (avoiding a situation in which a free model is released, but is used commercially)?
According to https://www.copyright.gov/fair-use/, using creative work is less likely to support a "fair use" claim; also, using a large amount of copyrighted work for derivative creation could play agains the "fair use" claim. Could the authors add a discussion on how the dataset addresses each of the requisites for fair use?
It would be interesting to have a professional opinion - and maybe even community guidelines in the future - for this matter, as the liability could be large; if that issue is solved, the dataset is amazing. Could the authors include a detailed legal analysis - maybe derived from this consultancy with a copyright expert - and include their findings in the paper, particularly addressing potential liabilities and jurisdictional differences (remember, we have a worldwide community, so these rules might be different in each country).