PT - JOURNAL ARTICLE AU - Van Poelvoorde, Laura A. E. AU - Delcourt, Thomas AU - Coucke, Wim AU - Herman, Philippe AU - De Keersmaecker, Sigrid C. J. AU - Saelens, Xavier AU - Roosens, Nancy AU - Vanneste, Kevin TI - Strategy and performance evaluation of low-frequency variant calling for SARS-CoV-2 in wastewater using targeted deep Illumina sequencing AID - 10.1101/2021.07.02.21259923 DP - 2021 Jan 01 TA - medRxiv PG - 2021.07.02.21259923 4099 - http://medrxiv.org/content/early/2021/07/07/2021.07.02.21259923.short 4100 - http://medrxiv.org/content/early/2021/07/07/2021.07.02.21259923.full AB - The ongoing COVID-19 pandemic, caused by SARS-CoV-2, constitutes a tremendous global health issue. Continuous monitoring of the virus has become a cornerstone to make rational decisions on implementing societal and sanitary measures to curtail the virus spread. Additionally, emerging SARS-CoV-2 variants have increased the need for genomic surveillance to detect particular strains because of their potentially increased transmissibility, pathogenicity and immune escape. Targeted SARS-CoV-2 sequencing of wastewater has been explored as an epidemiological surveillance method for the competent authorities. Few quality criteria are however available when sequencing wastewater samples, and those available typically only pertain to constructing the consensus genome sequence. Multiple variants circulating in the population can however be simultaneously present in wastewater samples. The performance, including detection and quantification of low-abundant variants, of whole genome sequencing (WGS) of SARS-CoV-2 in wastewater samples remains largely unknown. Here, we evaluated the detection and quantification of mutations present at low abundances using the SARS-CoV-2 lineage B.1.1.7 (alpha variant) defining mutations as a case study. Real sequencing data were in silico modified by introducing mutations of interest into raw wild-type sequencing data, or by mixing wild-type and mutant raw sequencing data, to mimic wastewater samples subjected to WGS using a tiling amplicon-based targeted metagenomics approach and Illumina sequencing. As anticipated, higher variation, lower sensitivity and more false negatives, were observed at lower coverages and allelic frequencies. We found that detection of all low-frequency variants at an abundance of 10%, 5%, 3% and 1%, requires at least a sequencing coverage of 250X, 500X, 1500X and 10,000X, respectively. Although increasing variability of estimated allelic frequencies at decreasing coverages and lower allelic frequencies was observed, its impact on reliable quantification was limited. This study provides a highly sensitive low-frequency variant detection approach, which is publicly available at https://galaxy.sciensano.be, and specific recommendations for minimum sequencing coverages to detect clade-defining mutations at specific allelic frequencies.Competing Interest StatementThe authors have declared no competing interest.Funding StatementThis study was financed by Sciensano through COVID-19 special fundingAuthor DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:No approval neededAll necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.YesAll data is publicly available.