Machine Learning in Clinical Journals: Moving from Inscrutable to Informative

Although machine learning (ML) algorithms have grown more prevalent in clinical journals, inconsistent reporting of methods has led to skepticism about ML results and has blunted their adoption into clinical practice. A common problem authors face when reporting on ML methods is the lack of a single reporting guideline that applies to the panoply of ML problem types and approaches. Responding to this concern, Stevens et al1 in this issue of Circulation: Cardiovascular Quality and Outcomes propose a set of reporting recommendations for ML papers. This is meant to augment existing (but more general) reporting guidelines on predictive models, such as the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) statement2— recognizing that international efforts are underway to address these concerns more fully, including TRIPOD-ML and Consolidated Standards of Reporting Trials–Artificial Intelligence (CONSORT-AI).3,4 However, a more fundamental question remains: will such guidelines address all of the different types of ML methods increasingly used in the published literature (such as unsupervised learning) based on their proposed scope? This is not a small problem in a dynamic field full of rapid change. Although Stevens et al focus on the spectrum of ML problem types between unsupervised and supervised learning, for instance, more recent literature has also seen the application of reinforcement learning methods to clinical problems.5,6 Additionally, the phrase “machine learning” can be used to refer to a number of different algorithms, which include decision trees, random forests, gradient-boosted machines, support vector machines, and neural networks. Perhaps somewhat confusingly, the phrase is also commonly applied to describe methods like linear regression and Bayesian models that readers of clinical journals are more familiar with. Thus, it is not uncommon to read an article purport to use “machine learning” in the title, only to find the actual algorithm is a penalized regression model, which a statistical reviewer may consider to be a traditional statistics model. As with many types of clinical research, we think standardized reporting of ML methods is a needed advancement and will be helpful for editors and reviewers evaluating the quality of manuscripts. Ultimately, this will be of great benefit to readers of the published article. In this editorial, however, we describe ongoing difficulties in determining what is and what is not an ML manuscript to which Stevens et al’s reporting recommendations would apply. We then discuss how this determination impacts how an article is perceived and judged by journals and readers. On the one hand, strong model performance described in articles may be the result of overfitting, and this could be prevented with better reporting practices. On the other hand, valid ML findings may be overlooked by even experienced statistiCirculation: Cardiovascular Quality and Outcomes

Paper

Full text

PDF

Machine Learning in Clinical Journals: Moving from Inscrutable to Informative

Semantic Scholar · Medicine · 2020

Abstract

Although machine learning (ML) algorithms have grown more prevalent in clinical journals, inconsistent reporting of methods has led to skepticism about ML results and has blunted their adoption into clinical practice. A common problem authors face when reporting on ML methods is the lack of a single reporting guideline that applies to the panoply of ML problem types and approaches. Responding to this concern, Stevens et al1 in this issue of Circulation: Cardiovascular Quality and Outcomes propose a set of reporting recommendations for ML papers. This is meant to augment existing (but more general) reporting guidelines on predictive models, such as the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis (TRIPOD) statement2— recognizing that international efforts are underway to address these concerns more fully, including TRIPOD-ML and Consolidated Standards of Reporting Trials–Artificial Intelligence (CONSORT-AI).3,4 However, a more fundamental question remains: will such guidelines address all of the different types of ML methods increasingly used in the published literature (such as unsupervised learning) based on their proposed scope? This is not a small problem in a dynamic field full of rapid change. Although Stevens et al focus on the spectrum of ML problem types between unsupervised and supervised learning, for instance, more recent literature has also seen the application of reinforcement learning methods to clinical problems.5,6 Additionally, the phrase “machine learning” can be used to refer to a number of different algorithms, which include decision trees, random forests, gradient-boosted machines, support vector machines, and neural networks. Perhaps somewhat confusingly, the phrase is also commonly applied to describe methods like linear regression and Bayesian models that readers of clinical journals are more familiar with. Thus, it is not uncommon to read an article purport to use “machine learning” in the title, only to find the actual algorithm is a penalized regression model, which a statistical reviewer may consider to be a traditional statistics model. As with many types of clinical research, we think standardized reporting of ML methods is a needed advancement and will be helpful for editors and reviewers evaluating the quality of manuscripts. Ultimately, this will be of great benefit to readers of the published article. In this editorial, however, we describe ongoing difficulties in determining what is and what is not an ML manuscript to which Stevens et al’s reporting recommendations would apply. We then discuss how this determination impacts how an article is perceived and judged by journals and readers. On the one hand, strong model performance described in articles may be the result of overfitting, and this could be prevented with better reporting practices. On the other hand, valid ML findings may be overlooked by even experienced statistiCirculation: Cardiovascular Quality and Outcomes

References (34)

Scroll for more · 22 remaining

Similar papers

© 2026 NYSGPT2525 LLC