PT - JOURNAL ARTICLE AU - Zhuang, Xiaowei AU - Vo, Van AU - Moshi, Michael A. AU - Dhede, Ketan AU - Ghani, Nabih AU - Akbar, Shahraiz AU - Chang, Ching-Lan AU - Young, Angelia K. AU - Buttery, Erin AU - Bendik, William AU - Zhang, Hong AU - Afzal, Salman AU - Moser, Duane AU - Cordes, Dietmar AU - Lockett, Cassius AU - Gerrity, Daniel AU - Kan, Horng-Yuan AU - Oh, Edwin C. TI - Early Detection of Novel SARS-CoV-2 Variants from Urban and Rural Wastewater through Genome Sequencing and Machine Learning AID - 10.1101/2024.04.18.24306052 DP - 2024 Jan 01 TA - medRxiv PG - 2024.04.18.24306052 4099 - http://medrxiv.org/content/early/2024/04/19/2024.04.18.24306052.short 4100 - http://medrxiv.org/content/early/2024/04/19/2024.04.18.24306052.full AB - Genome sequencing from wastewater has emerged as an accurate and cost-effective tool for identifying SARS-CoV-2 variants. However, existing methods for analyzing wastewater sequencing data are not designed to detect novel variants that have not been characterized in humans. Here, we present an unsupervised learning approach that clusters co-varying and time-evolving mutation patterns leading to the identification of SARS-CoV-2 variants. To build our model, we sequenced 3,659 wastewater samples collected over a span of more than two years from urban and rural locations in Southern Nevada. We then developed a multivariate independent component analysis (ICA)-based pipeline to transform mutation frequencies into independent sources with co-varying and time-evolving patterns and compared variant predictions to >5,000 SARS-CoV-2 clinical genomes isolated from Nevadans. Using the source patterns as data-driven reference “barcodes”, we demonstrated the model’s accuracy by successfully detecting the Delta variant in late 2021, Omicron variants in 2022, and emerging recombinant XBB variants in 2023. Our approach revealed the spatial and temporal dynamics of variants in both urban and rural regions; achieved earlier detection of most variants compared to other computational tools; and uncovered unique co-varying mutation patterns not associated with any known variant. The multivariate nature of our pipeline boosts statistical power and can support accurate and early detection of SARS-CoV-2 variants. This feature offers a unique opportunity for novel variant and pathogen detection, even in the absence of clinical testing.Competing Interest StatementThe authors have declared no competing interest.Funding StatementVV, ECO are supported by NIH grants: GM103440 and MH109706 and a CARES Act grant from the Nevada Governor's Office of Economic Development. VV, CL, DG, HK, and ECO are supported by Grant Number NH75OT000057-01-00 from the Centers for Disease Control and Prevention. The project contents are solely the responsibility of the authors and do not necessarily represent the official views of the Centers for Disease Control and Prevention. DPM was supported by the Nevada Water Resources Research Institute/USGS under Grant/Cooperative Agreement No. G21AP10578 through the Division of Hydrologic Sciences at Desert Research Institute.Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesI confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesAll data produced in the present study are available upon reasonable request to the authors