ABSTRACT
We present a novel methodology for subtyping of persons with a common clinical symptom complex by integrating heterogeneous continuous and categorical data. We illustrate it by clustering women with lower urinary tract symptoms (LUTS), who represent a heterogeneous cohort with overlapping symptoms and multifactorial etiology. Identifying subtypes within this group would potentially lead to better diagnosis and treatment decision-making. Data collected in the Symptoms of Lower Urinary Tract Dysfunction Research Network (LURN), a multi-center prospective observational cohort study, included self-reported urinary and non-urinary symptoms, bladder diaries, and physical examination data for 545 women. Heterogeneity in these multidimensional data required thorough and non-trivial preprocessing, including scaling by controls and weighting to mitigate data redundancy, while the various data types (continuous and categorical) required novel methodology using a weighted Tanimoto indices approach. Data domains only available on a subset of the cohort were integrated using a semi-supervised clustering approach. Novel contrast criterion for determination of the optimal number of clusters in consensus clustering was introduced and compared with existing criteria. Distinctiveness of the clusters was confirmed by using multiple criteria for cluster quality, and by testing for significantly different variables in pairwise comparisons of the clusters. Cluster dynamics were explored by analyzing longitudinal data at 3- and 12-month follow-up. Five distinct clusters of women with LUTS were identified using the developed methodology. The clinical relevance of the identified clusters is discussed and compared with the current conventional approaches to the evaluation of LUTS patients. Rationale and thought process are described for selection of procedures for data preprocessing, clustering, and cluster evaluation. Suggestions are provided for minimum reporting requirements in publications utilizing clustering methodology with multiple heterogeneous data domains.
Competing Interest Statement
The authors have declared no competing interest.
Clinical Trial
NCT02485808
Funding Statement
This study is supported by the National Institute of Diabetes & Digestive & Kidney Diseases through cooperative agreements. Grant Numbers: DK097780 DK097772 DK097779 DK099932 DK100011 DK100017 DK099879. Dr. Andreev Biomarker Ancillary LURN R01 Grant Number: 5R01DK125251. Research reported in this publication was supported at Northwestern University, in part, by the National Institutes of Health National Center for Advancing Translational Sciences. Grant Number: UL1TR001422. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
The authors confirm all relevant ethical guidelines have been followed, and all research has been conducted according to the principles expressed in the Declaration of Helsinki. Informed consent has been obtained from participants. Institutional Review Board (IRB) approval has been obtained from: Ethical and Independent Review Services (E&I) IRB, an Association for the Accreditation of Human Research Protection Programs (AAHRPP) Accredited Board, Registration #IRB 00007807.
All necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.
Yes
Data Availability
The data that support the findings of this study are openly available in the NIDDK Central Repository at https://repository.niddk.nih.gov/; please reference the acronym LURN.