PT - JOURNAL ARTICLE AU - Barmpas, Petros AU - Tasoulis, Sotiris AU - Vrahatis, Aristidis G. AU - Anagnostou, Panagiotis AU - Georgakopoulos, Spiros AU - Prina, Matthew AU - Ayuso-Mateos, José Luis AU - Bickenbach, Jerome AU - Bayes, Ivet AU - Bobak, Martin AU - Caballero, Francisco Félix AU - Chatterji, Somnath AU - Egea-Cortés, Laia AU - García-Esquinas, Esther AU - Leonardi, Matilde AU - Koskinen, Seppo AU - Koupil, Ilona AU - Pająk, Andrzej AU - Prince, Martin AU - Sanderson, Warren AU - Scherbov, Sergei AU - Tamosiunas, Abdonas AU - Galas, Aleksander AU - MariaHaro, Josep AU - Sanchez-Niubo, Albert AU - Plagianakos, Vassilis P. AU - Panagiotakos, Demosthenes TI - Unsupervised Learning for Large Scale Data: The ATHLOS Project AID - 10.1101/2021.04.01.21254751 DP - 2021 Jan 01 TA - medRxiv PG - 2021.04.01.21254751 4099 - http://medrxiv.org/content/early/2021/04/06/2021.04.01.21254751.short 4100 - http://medrxiv.org/content/early/2021/04/06/2021.04.01.21254751.full AB - Recent technological advancements in various domains, such as the biomedical and health, offer a plethora of big data for analysis. Part of this data pool is the experimental studies that record various and several features for each instance. It creates datasets having very high dimensionality with mixed data types, with both numerical and categorical variables. On the other hand, unsupervised learning has shown to be able to assist in high-dimensional data, allowing the discovery of unknown patterns through clustering, visualization, dimensionality reduction, and in some cases, their combination. This work highlights unsupervised learning methodologies for large-scale, high-dimensional data, providing the potential of a unified framework that combines the knowledge retrieved from clustering and visualization. The main purpose is to uncover hidden patterns in a high-dimensional mixed dataset, which we achieve through our application in a complex, real-world dataset. The experimental analysis indicates the existence of notable information exposing the usefulness of the utilized methodological framework for similar high-dimensional and mixed, real-world applications.Competing Interest StatementThe authors have declared no competing interest.Funding StatementThis work is supported by the ATHLOS (Aging Trajectories of Health: Longitudinal Opportunities and Synergies) project, funded by the European Union's Horizon 2020 Research and Innovation Program under grant agreement number 635316.Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:Does not apply in our workAll necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.YesData sharing is not applicable to this article