Back

Fair molecular feature selection unveils universally tumor lineage-informative methylation sites in colorectal cancer

Li, X. C.; Liu, Y.; Schaffer, A. A.; Mount, S.; Sahinalp, S. C.

2024-02-27 bioinformatics
10.1101/2024.02.22.580595 bioRxiv
Show abstract

In the era of precision medicine, performing comparative analysis over diverse patient populations is a fundamental step towards tailoring healthcare interventions. However, the critical aspect of equitably selecting molecular features across multiple patients is often overlooked. To address this challenge, we introduce FALAFL (FAir muLti-sAmple Feature seLection), an algorithmic approach based on combinatorial optimization. FALAFL is designed to bridge the gap between molecular feature selection and algorithmic fairness, ensuring a fair selection of molecular features from all patient samples in a cohort. We have applied FALAFL to the problem of selecting lineage-informative CpG sites within a cohort of colorectal cancer patients subjected to low read coverage single-cell methylation sequencing. Our results demonstrate that FALAFL can rapidly and robustly determine the optimal set of CpG sites, which are each well covered by cells across the vast majority of the patients, while ensuring that in each patient a high proportion of these sites have good read coverage. An analysis of the FALAFL-selected sites reveals that their tumor lineage-informativeness exhibits a strong correlation across a spectrum of diverse patient profiles. Furthermore, these universally lineage-informative sites are highly enriched in the inter CpG island regions. FALAFL integrates equity considerations into the molecular feature selection from single-cell sequencing data obtained from a patient cohort. We hope that it will help propel equitable healthcare data science practices and contribute to the advancement of our understanding of complex diseases.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.