Machine Learning for Electrode Materials: Property Prediction via Composition

In this work, we benchmark three leading composition based Machine Learning (ML) frameworks, MODNet, CrabNet, and a random forest model based on Magpie features, predicting the properties of battery electrode materials using the Materials Project Battery Explorer dataset. We evaluate these models based on predictive accuracy, visualize numerical features using two-dimensional embeddings, and quantify performance using standard metrics. Our results demonstrate that CrabNet consistently outperforms the other models across all tests. To validate these findings, we employ bootstrap resampling and two cross-validation (CV) strategies (leave-one-cluster-out and stratified 5-fold CV), comparing each model against a control baseline, using unseen experimental data as a hold-out test. We also apply unsupervised clustering using t-SNE and DBSCAN on physically observed features extracted from matminer, revealing coherent material groupings without prior labels. The final selected model consistently improves over controls, and we believe can be used as a early stage oracle for electrode materials composition screening. Our study aims to identify the error distributions and limitations of the approach, discussing the challenges with developing robust ML models. Despite these constraints, our findings suggest the final selected model is effective for early-stage compositional screening.

Paper

Similar papers

© 2026 NYSGPT2525 LLC