Abstract
Earlier detection of pancreatic ductal adenocarcinoma (PDAC) is key to improving patient outcomes, as it is mostly detected at advanced stages which are associated with poor survival. Developing non-invasive blood tests for early detection would be an important breakthrough. The primary objective of the work presented here was to use a unique dataset, that is both large and prospectively collected, to quantify a set of 96 cancer-associated proteins and construct multi-marker models with the capacity to accurately predict PDAC years before diagnosis. The data is part of a nested case control study within UK Collaborative Trial of Ovarian Cancer Screening and is comprised of 219 samples, collected from a total of 143 post-menopausal women who were diagnosed with pancreatic cancer within 70 months after sample collection, and 248 matched non-cancer controls. We developed a stacked ensemble modelling technique to achieve robustness in predictions and, therefore, improve performance in newly collected datasets. With a pool of 10 base-learners and a Bayesian averaging meta-learner, we can predict PDAC status with an AUC of 0.91 (95% CI 0.75 - 1.0), sensitivity of 92% (95% CI 0.54 - 1.0) at 90% specificity, up to 1 year to diagnosis, and at an AUC of 0.85 (95% CI 0.74 - 0.93) up to 2 years to diagnosis (sensitivity of 61%, 95 % CI 0.17 - 0.83, at 90% specificity). These models also use clinical covariates such as hormone replacement therapy use (at randomization), oral contraceptive pill use (ever) and diabetes and outperform biomarker combinations cited in the literature.
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This research was funded by Cancer Research UK (grant C12077/A26223) and supported by the Pancreatic Cancer UK Early Diagnosis Award 2018, project The Accelerated Diagnosis of neuroEndocrine and Pancreatic TumourS (ADEPTS), and by the National Institute for Health Research (NIHR) University College London Biomedical Research Centre. UKCTOCS was core funded by the Medical Research Council, Cancer Research UK, and the Department of Health with additional support from the Eve Appeal, Special Trustees of Bart's and the London, and Special Trustees of UCLH. AZ thanks support from the Ministry of Science and Higher Education of the Russian Federation within the framework of state support for the creation and development of World-Class Research Centers Digital biodesign and personalized healthcare Number 075-15-2020-926. HW is supported by the NIHR Imperial Biomedical Research Centre. EC is supported by Cancer Research UK (C7690/A26881).
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
This nested case control discovery study was approved by the Joint UCL/UCLH Research Ethics Committee A (Ref. 05/Q0505/57).
I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.
Yes
Footnotes
The name of the platform used to screen 92 of the proteins used in this paper has been corrected. Table S1 has also been corrected.
Data Availability
Data requestors will need to sign a data access agreement and in keeping with patient consent for secondary use, obtain ethical approval for any new analyses. All packages used in the pre-processing of data and subsequent analysis have been identified to secure full reproducibility.