Deep learning for blood glucose prediction has advanced quickly, and these models are intended for safety-critical settings such as artificial pancreas systems. Yet little is known about whether published results can actually be reproduced, or whether they hold up on populations other than the one they were developed on.
What we did
We reviewed 67 recent papers proposing deep learning methods for glucose prediction to identify recurring reproducibility obstacles. We then reimplemented eight representative methods and evaluated them under a standardized framework covering technical, statistical, and conceptual reproducibility — using over 1.36 million CGM samples (5,061 days) from 128 individuals with type 1 diabetes across three public datasets: OhioT1DM, DiaTrend, and T1DEXI.
What we found
- The models were largely technically and statistically reproducible — reimplementations behaved consistently on their original data.
- Conceptual reproducibility was limited: performance degraded when models were moved to datasets reflecting different diabetes management patterns.
- Prediction error was strongly tied to individual glycemic control. Participants who spent less time in the target range (70–180 mg/dL) were consistently predicted worst — precisely the people such systems most need to serve.
These results argue for greater transparency, more diverse datasets, standardized evaluation practices, and accessible code before these models can be relied upon clinically (Lu et al., 2026).
References
2026
-
Deep learning for blood glucose prediction: Reproducibility challenges and factors affecting differential performance
Baiying Lu, Biratal Wagle, Zhaohui Liang, and 2 more authors
PLOS Digital Health, Sep 2026
Blood glucose prediction is a critical component of next-generation diabetes technologies, such as artificial pancreas systems, where reliable performance is essential for safety and effectiveness. Although deep learning methods have achieved promising advances in this area, a critical gap remains in understanding the reproducibility and generalizability of these methods. To contextualize the gap, this study reviewed 67 recent papers that proposed a deep learning method for glucose prediction to identify key reproducibility challenges. Next, we adopted a standardized framework, encompassing technical, statistical, and conceptual reproducibility evaluations, to experimentally assess the reproducibility of eight representative deep learning methods. To achieve this, we reimplemented and evaluated these eight deep learning methods using over 1.36 million continuous glucose monitoring samples (5,061 days) from 128 individuals with type 1 diabetes across three public datasets: OhioT1DM, DiaTrend, and T1DEXI. We found that even though these models demonstrated good technical and statistical reproducibility, their conceptual reproducibility—the ability to generalize to datasets with different diabetes management patterns—was limited. Further analyses revealed that each model’s overall prediction performance was strongly influenced by individual glycemic control, with higher prediction errors observed among participants with lower time with blood glucose in the target range (70–180 mg/dL). This study identified key reproducibility challenges associated with current blood glucose prediction methods within type 1 diabetes populations, highlighting the need for increased transparency, dataset diversity, standardized evaluation practices, and code accessibility to ensure reproducible and reliable models for blood glucose prediction.