Personalized Forecasting of Glycemic Control in Type 1 and 2 Diabetes Using Foundational AI and Machine Learning Models
Accurate week-ahead forecasts of continuous glucose monitoring (CGM)–derived metrics could enable proactive diabetes management, but relative performance of modern tabular learning approaches is incompletely defined. We trained and internally validated four regression models (CatBoost, XGBoost, AutoGluon, tabPFN) to predict 6-week-ahead CGM metrics (time-in-range [TIR], time-in-tight-range [TITR], time-above-range [TAR], time-below-range [TBR], coefficient of variation, Mean Amplitude of Glycemic Excursions [MAGE], and related quantiles) using 4622 case-weeks from two cohorts (T1DM, n = 3389; T2DM, n = 1233). Performance was assessed with mean absolute error (MAE) and mean absolute relative difference (MARD); quantile classification was summarized via confusion-matrix heatmaps. Across T1DM and T2DM, all models produced broadly comparable performance for most targets. For T1DM, MARD for TIR, TITR, TAR, and MAGE ranged 8.5%–16.5% while TBR showed large MARD (mean ≈48%) despite low MAE. AutoGluon and tabPFN showed lower MAE than XGBoost for several targets (e.g., TITR: P < 0.01; TAR/TBR: P < 0.05–0.01). For T2DM, MARD ranged 7.8%–23.9%, and TBR relative error was ≈78%; tabPFN outperformed other models for TIR ( P < 0.01), and AutoGluon/tabPFN outperformed CatBoost/XGBoost on TAR ( P < 0.05). Inference time per 1000 cases varied markedly (PFN 699 s; AG 2.7 s; CatBoost 0.04 s, XGBoost 0.04 s). Week-ahead CGM metrics are predictable with reasonable accuracy using modern tabular models, but low-prevalence hypoglycemia remains difficult to predict in relative terms. Advanced automated machine learning and foundation models yield modest accuracy gains at substantially higher computational cost. External validation is required before these tools can be considered ready for clinical implementation.