Back

Fast Loading and Transformation of Large-Volume Bioimaging Data Consistently Approaching the I/O Peak Speed

Cai, L.; Qu, X.; Zhou, H.; Li, N.; Gou, X.; Li, S.; Huang, J.; Guo, S.; Liu, X.; Lv, X.; Quan, T.; Zeng, S.

2025-12-12 bioinformatics
10.64898/2025.12.10.693603 bioRxiv
Show abstract

High-resolution large-volume biological imaging techniques are now widely used in biological research, but inefficiencies in data processing and visualization persist due to bottlenecks in loading/saving pipelines and limitations of conventional pyramid formats. To address these issues, we have decoupled data loading process into reading from disk and in-memory decompression, while data saving has been separated into in-memory compression and writing to disk. This approach allows us to dedicate more CPU cores to decompression and compression tasks, optimizing resource allocation and maximizing disk I/O throughput. As a result, both loading and saving processes operate near the disks peak data transfer capacity. Further, we integrated a video stream-inspired format into the pyramid structure, allowing direct access to regions of interest (ROIs) without loading extraneous data into RAM, reducing memory overhead. Compared to the latest methods, our approach improves pyramid data transformation efficiency by at least 6x and dramatically accelerates downstream tasks like visualization and deep learning inference. TeaserNovel bioimaging data processing pipeline achieves near-peak I/O throughput via decoupled architecture and video-formatting.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.