Back

Bio-medical Big Data Operating System (Bio-OS): An Integrated Data Mining Environment for Data Intensive Scientific Research

Liu, J.; Yang, L.; Xiao, Q.; Li, Z.; Chen, J.; Chen, Z.; Zhou, J.; Wan, X.; Tsung-Hui Chang, T.-H.; Zhang, X.; Li, Y.

2024-10-20 bioinformatics
10.1101/2024.10.17.612997 bioRxiv
Show abstract

The advent of high throughput sequencing has ushered life science and clinical research into the era of big data, posing significant challenges for reproducibility due to the complexity of data integration and analysis. Although the FAIR principles advocate for the transparent and reliable sharing of scientific data, their implementation remains hampered by technical barriers. The Global Alliance for Genomics and Health (GA4GH) has made strides in standardizing data and tools, yet a comprehensive solution for reproducibility is lacking. In response, we present BioOS, an open source, cloud native Biomedical big data Operating System. This system encapsulates study components data, code, tools, and environments into workspaces, enhancing reproducibility and validation. BioOS employs JSON Schema for machine readability and includes a Hierarchy Hash Mechanism to ensure data integrity. Adhering to GA4GH protocols, BioOS simplifies complex technological implementations, making advanced research tools accessible. Demonstrated through representative workspaces, BioOS fosters seamless research replication, peer review, and editorial evaluation. Its cloud native infrastructure supports dynamic resource allocation, enabling efficient handling of large scale analyses. By integrating AI driven Large Language Models, BioOS enhances user interaction and operational flexibility. As an evolving open source platform, BioOS exemplifies a transformative approach to biomedical research, aligning with FAIR principles and advancing the AI for Science paradigm, thus promoting a more connected, efficient, and impactful research environment.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.