Abstract
Introduction Existing methods to make data Findable, Accessible, Interoperable, and Reusable (FAIR) are usually carried out in a post-hoc manner: after the research project is conducted and data are collected. De-novo FAIRification, on the other hand, incorporates the FAIRification steps in the process of a research project. In medical research, data is often collected and stored via electronic Case Report Forms (eCRFs) in Electronic Data Capture (EDC) systems. By implementing a de-novo FAIRification process in such a system, the reusability and, thus, scalability of FAIRification across research projects can be greatly improved. In this study, we developed and implemented a novel method for de-novo FAIRification via an EDC system. We evaluated our method by applying it to the Registry of Vascular Anomalies (VASCA).
Methods Our EDC and research project independent method ensures that eCRF data entered into an EDC system can be transformed into machine-readable, FAIR data using a semantic data model (a canonical representation of the data, based on ontology concepts and semantic web standards) and mappings from the model to questions on the eCRF. The FAIRified data are stored in a triple store and can, together with associated metadata, be accessed and queried through a FAIR Data Point. The method was implemented in Castor EDC, an EDC system, through a data transformation application. The FAIRness of the output of the method, the FAIRified data and metadata, was evaluated using the FAIR Evaluation Services.
Results We successfully applied our FAIRification method to the VASCA registry. Data entered on eCRFs is automatically transformed into machine-readable data and can be accessed and queried using SPARQL queries in the FAIR Data Point. Twenty-one FAIR Evaluator tests pass and one test regarding the metadata persistence policy fails, since this policy is not in place yet.
Conclusion In this study, we developed a novel method for de-novo FAIRification via an EDC system. Its application in the VASCA registry and the automated FAIR evaluation show that the method can be used to make clinical research data FAIR when they are entered in an eCRF without any intervention from data management and data entry personnel. Due to the generic approach and developed tooling, we believe that our method can be used in other registries and clinical trials as well.
Competing Interest Statement
MK is employed by Castor, the Electronic Data Capture platform that was used for data collection. DA is Castor’s CEO. The remaining authors state no conflicts of interest.
Funding Statement
MK’s and DA’s work is supported by funding from Castor. AJ, BV, RK, PAC’tH, RC and MR’s work is supported by the funding from the European Union’s Horizon 2020 research and innovation programme under the EJP RD COFUND-EJP No 825575. BV and LSK are members of the Vascular Anomalies Working Group (VASCA WG) of the European Reference Network for Rare Multisystemic Vascular Diseases (VASCERN) - Project ID: 769036. KG’s work is supported by the department of Medical Imaging, Radboud University Medical Center.
Author Declarations
I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.
Yes
The details of the IRB/oversight body that provided approval or exemption for the research described are given below:
Not applicable
All necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived.
Yes
I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).
Yes
I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable.
Yes
Data Availability
Not applicable
Abbreviations
- AAI
- Authentication and Authorization Infrastructures
- API
- Application Programming Interface
- CDE
- Common Data Element
- DC
- Dublin Core
- DCAT2
- Data Catalogue Vocabulary version 2
- eCRF
- electronic Case Report Form
- EDC
- Electronic Data Capture
- ERN
- European Reference Network
- EU
- European Union
- FAIR
- Findable, Accessible, Interoperable, and Reusable
- JRC
- Joint Research Center
- NIH
- National Institute of Health
- NWO
- The Dutch Research Council
- RD
- rare disease
- RDF
- Resource Description Framework
- VASCA
- Vascular Anomalies
- VASCERN
- European Reference Network on rare vascular diseases