📄 Sciences Methods and Technologies
International Journal (SciMeTech)

Volume 2 · Issue 1 · 2026
ISSN: 3085-5284
A Comprehensive Review of Missing Data Handling: From Traditional Methods to Modern AI-driven Solutions
Zakaria JANAH, Anas ABOU EL KALAM, Wissam AABASS
Pages 1–10 · Cadi Ayyad University, National School of Applied Sciences, LaRTID Laboratory, Marrakech, Morocco
Abstract
This is the problem of incomplete datasets, which is highly common and persistent in all spheres of empirical inquiry and casts a fundamental challenge on the quality of statistical inferences and predictive capabilities of machine learning models. This review gives a detailed analysis of the development of techniques that are applied to manage missing data, their primitive methods of historical analysis to modern and sophisticated techniques. Our discussion begins with a description of the necessary typology of missing data mechanisms, Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR), as it offers the theory behind selecting an appropriate imputation strategy. The paper subsequently follows up the history of the traditional techniques of statistics, showing the inadequacy of such simplistic approaches to the data as case deletion and single imputation in terms of the resilient, uncertainty-based Multiple Imputation (MI), which continues to be considered the standard of MAR data. Next, we discuss the new paradigm of AI and ML-based models, including Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and Transformers that have shown themselves to be highly effective at representing complex non-linear patterns in complex high-dimensional data. There is a particular emphasis on practical challenges of handling the heterogeneous types of data. The analysis we do in comparison reveals that no universal solution is universally best, and the choice will have to be informed by the type of the missingness, as well as the arrangement of the data, and the analytic goals. The review ends with the statement that even though modern AI models can make spectacular predictions, much work is required to establish reliable approaches to MNAR settings, formalize uncertainty estimation in deep learning imputations, and create standardized evaluation methods of methods.
Keywords: Missing Data, Imputation, Multiple Imputation, Machine Learning, Deep Learning

References

  1. Holmes, J., Sacchi, L., & Bellazzi, R. (2004). Artificial intelligence in medicine. Ann R Coll Surg Engl, 86, 334-8.
  2. Rubin, D. B. (1976). Inference and missing data. Biometrika, 63(3), 581-592.
  3. Little, R. J. A., & Rubin, D. B. (2019). Statistical analysis with missing data (3rd ed.). John Wiley & Sons.
  4. National Research Council. (2010). Drawing inferences from incomplete data. In The prevention and treatment of missing data in clinical trials. Washington, DC: The National Academies Press.
  5. Kang, H. (2013). The prevention and handling of the missing data. Korean J. Anesthesiol., 64(5), 402–406. https://doi.org/10.4097/kjae.2013.64.5.402
  6. Chourib, I. (2025). Missing data handling: A comprehensive review, taxonomy, and comparative evaluation. J. Comput. Commun., 13(5), 1–28. https://doi.org/10.4236/jcc.2025.135001
  7. Heymans, M. W., & Twisk, J. W. R. (2022). Handling missing data in clinical research. J. Clin. Epidemiol., 151, 185–188. https://doi.org/10.1016/j.jclinepi.2022.08.016
  8. Raghunathan, T. E., Lepkowski, J. M., & van Buuren, S. (2018). Multiple imputation: A review of practical and theoretical findings. Stat. Sci., 33(2), 211–232. https://doi.org/10.1214/18-STS644
  9. van Buuren, S. (2018). Flexible imputation of missing data (2nd ed.). Chapman & Hall/CRC.
  10. Wang, Z., Tyshetskiy, Y., Lee, D.-H., & Compton, P. (2024). Deep learning for multivariate time series imputation: A survey. arXiv preprint arXiv:2402.04059. https://doi.org/10.48550/arXiv.2402.04059
  11. Zhong, Y., & Gao, L. (2023). Deep learning methods for omics data imputation. Genes, 14(10), 1963. https://doi.org/10.3390/genes14101963
  12. Zhou, Y., Aryal, S., & Bouadjenek, M. R. (2024). Review for Handling Missing Data with special missing mechanism. arXiv. https://doi.org/10.48550/ARXIV.2404.04905
  13. He, Y., Zaslavsky, A. M., Harrington, D. P., Catalano, P., & Landrum, M. B. (2010). Multiple imputation in a large-scale complex survey: A practical guide. Stat. Med., 29(29), 3003–3016. https://doi.org/10.1002/sim.4034
  14. van Buuren, S., Brand, J. P. L., Groothuis-Oudshoorn, C. G. M., & Rubin, D. B. (2006). Fully conditional specification in multivariate imputation. J. Stat. Comput. Simul., 76(12), 1049–1064. https://doi.org/10.1080/00949650600956996
  15. Joel, L. O., Doorsamy, W., & Paul, B. S. (2025). A comparative study of imputation techniques for missing values in healthcare diagnostic datasets. Int. J. Data Sci. Anal. https://doi.org/10.1007/s41060-025-00825-9
  16. Rácz, A., & Gere, A. (2025). Comparison of missing value imputation tools for machine learning models based on product development cases studies. LWT, 221, 117585. https://doi.org/10.1016/j.lwt.2025.117585