Abstract
Abstract
Missing data are common in environmental mixture studies and can bias inference if not properly addressed. This study evaluates how different imputation strategies influence variable selection performance in six mixture modeling frameworks: Weighted Quantile Sum (WQS), Bayesian WQS (BWQS), Quantile g-computation (Q-gcomp), Bayesian Kernel Machine Regression (BKMR), Elastic Net, and Least Absolute Shrinkage and Selection Operator (LASSO). Using Monte Carlo simulations (MCS), we generated multivariate normal (MVNORM) and multivariate t- (MVT) distributed exposures under linear and nonlinear outcome structures. Missingness was introduced at 25% for exposure and outcome variables, and at 5%–25% under both missing-at-random (MAR) and missing-not-at-random (MNAR) mechanisms. The study compares single imputation (SI) methods (mean, median, and k-nearest neighbors [KNN]) with multiple imputation approaches [MI] (MICE and Amelia) based on sensitivity (SE), specificity (SP), and false discovery rate (FDR). Across simulations, MI consistently improved variable selection performance under MAR, whereas SI and listwise deletion increased variability and reduced accuracy. Under MNAR, performance declined across all methods, with greater instability observed in flexible models such as BKMR and WQS. Q-gcomp demonstrated the most consistent balance across SE, SP, and FDR and remained relatively robust to violations of the MAR assumption. Additional analyses highlight the importance of imputation model specification: excluding relevant covariates (e.g. income) from the imputation process increased variability and reduced stability, particularly for flexible models and binary outcomes. Application to NHANES 2007–2014 data (n = 8,233) showed that MI improved the stability of variable importance estimates, with BWQS and Q-gcomp yielding more reproducible exposure rankings. Overall, results demonstrate that both the missingness mechanism and the choice and specification of imputation strategy substantially influence variable selection in mixture models, underscoring the importance of carefully designed imputation procedures in environmental epidemiology.
Direct answer
What can I do from this paper page?
Use this page to scan "Evaluating Missing Data Imputation Strategies for Environmental Mixture Models: A Simulation Study and Applied Analysis" quickly: start with the summary and abstract, then check the authors, source, topics, and related papers. From here, open Scollr to follow Bayesian Methods and Mixture Models research, save the paper, or map adjacent work.
Research areas
Follow related topics
Citation
BibTeX
@article{Boafo2026Evaluating,
title = {Evaluating Missing Data Imputation Strategies for Environmental Mixture Models: A Simulation Study and Applied Analysis},
author = {Yvonne S. Boafo and Sayed Mostafa and Emmanuel Obeng-Gyasi},
journal = {Data Science in Science},
year = {2026},
doi = {10.1080/26941899.2026.2696661},
url = {https://doi.org/10.1080/26941899.2026.2696661}
}
FAQ
Using this paper in a discovery workflow
How do I find related work for this paper?
Use the related papers and topic links on this page as starting points. In Scollr, you can also open the paper and build a literature map around its references, citing papers, and related work.
How can I keep up with new Bayesian Methods and Mixture Models research papers?
Follow Bayesian Methods and Mixture Models research in Scollr. New papers from the topic flow into a personalized feed, and you can save useful studies to revisit later.
Can I cite this paper from this page?
This page includes a static BibTeX block for Evaluating Missing Data Imputation Strategies for Environmental Mixture Models: A Simulation Study and Applied Analysis. Always verify the DOI, source, and publication details against the publisher record before submitting a manuscript.
Follow this research in Scollr
Follow the topics and authors behind this paper, save useful studies, and build a literature map when you are ready to go deeper.
Get the app