Back

A Multi-omics Data Analysis Workflow Packaged as a FAIR Digital Object

Niehues, A.; de Visser, C.; Hagenbeek, F. A.; Kulkarni, P.; Pool, R.; Karu, N.; Kindt, A. S. D.; Singh, G.; Vermeiren, R. R. J. M.; Boomsma, D. I.; van Dongen, J.; 't Hoen, P. A. C.; van Gool, A. J.

2023-06-09 bioinformatics
10.1101/2023.06.07.543986 bioRxiv
Show abstract

BackgroundApplying good data management and FAIR data principles (Findable, Accessible, Interoperable, and Reusable) in research projects can help disentangle knowledge discovery, study result reproducibility, and data reuse in future studies. Based on the concepts of the original FAIR principles for research data, FAIR principles for research software were recently proposed. FAIR Digital Objects enable discovery and reuse of Research Objects, including computational workflows for both humans and machines. Practical examples can help promote the adoption of FAIR practices for computational workflows in the research community. We developed a multi-omics data analysis workflow implementing FAIR practices to share it as a FAIR Digital Object. FindingsWe conducted a case study investigating shared patterns between multi-omics data and childhood externalizing behavior. The analysis workflow was implemented as a modular pipeline in the workflow manager Nextflow, including containers with software dependencies. We adhered to software development practices like version control, documentation, and licensing. Finally, the workflow was described with rich semantic metadata, packaged as a Research Object Crate, and shared via WorkflowHub. ConclusionsAlong with the packaged multi-omics data analysis workflow, we share our experiences adopting various FAIR practices and creating a FAIR Digital Object. We hope our experiences can help other researchers who develop omics data analysis workflows to turn FAIR principles into practice. O_TEXTBOXKey PointsO_LIThe FAIR4RS principles provide guidelines to enhance the discovery and reuse of research software. C_LIO_LIFAIR Digital Objects support Findability, Accessibility, Interoperability, and Reusability by both humans and machines. C_LIO_LIWe here demonstrate the implementation multi-omics data analysis workflow and share it as a FAIR Digital Object. C_LI C_TEXTBOX

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.