Abstract
West Nile virus disease is a growing issue with devastating outbreaks and linkage to climate. It’s a complex disease with many factors contributing to emergence and spread. High-performance machine learning models, such as XGBoost, hold potential for development of predictive models which performs well with complex diseases like West Nile virus disease. Such models furthermore allow for expanded ability to discover biological, ecological, social and clinical associations as well as interaction effects. In 1951, a deductive method based on cooperative game theory was introduced: Shapley values. The Shapley method has since been shown to be the only way to derive “true” effect estimations from complex systems. Up till recently, however, wide-scale application has been computationally prohibitive. Herein, we present a novel implementation of the Shapley method applied to machine learning to derive high-quality effect estimations. We set out to apply this method to study the drivers of and predict West Nile virus in Europe. Model validity was furthermore tested using observed information in the time periods following the prospective prediction window. We furthermore benchmarked results of XGBoost models against equivalently specified logistic regression models. High predictive performance was consistently observed. All models were statistically equivalent in terms of AUC performance (96.3% average). The top features across models were found to be vapor pressure, the autoregressive past year’s feature, maximum temperature, wind speed, and local GNP. Moreover, when aggregated across quarters, we found that the effect of these features are broadly consistent across model configurations. We furthermore confirmed that for an equivalent level of model sophistication, XGBoost and logistic regressions performed similarly, with an advantage to XGBoost as model complexity increased. Our findings highlight the importance of ecological factors, such as climate, in determining outbreak risk of West Nile virus in Europe. We conclude by demonstrating the feasibility of same-year prospective early warning models that combine same-year observed climate with autoregressive geospatial covariates and long-term bioclimatic features. Scenario-based forecasts could likely be developed using similar methods, to provide for long-term intervention and resource planning, therefore increasing public health preparedness and resilience.
HighlightsFor geospatial analysis, XGBoost’s high-powered predictions are not always empirically sound
SHAP, an AI-driven enhancement to XGBoost, resolves this issue by: 1) deriving empirically-valid models for each individual case-region, and 2) setting classification thresholds accordingly
SHAP therefore allows for predictive consistency across models and improved generalizeability
Aggregate effect estimations produced by SHAP are consistent across model configurations
AI-driven methods improve model validity with respect to predicted range and determinants
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
Funding was received from the European Center for Disease Control and Prevention (contract no ECDC.9504).
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
n/a
All necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.
Yes
The Chan Zuckerberg Initiative, Cold Spring Harbor Laboratory, the Sergey Brin Family Foundation, California Institute of Technology, Centre National de la Recherche Scientifique, Fred Hutchinson Cancer Center, Imperial College London, Massachusetts Institute of Technology, Stanford University, University of Washington, and Vrije Universiteit Amsterdam.