Back

Long-read nanopore sequencing uncovers population-specific structural variation in the Middle East and North Africa

AL Yazeedi, T.; Tandonnet, S.; Hauns, S.; Almansoori, S.; Davis, P.; Backofen, R.; Tayoun, A. A.; Alkhnbashi, O.

2026-02-23 genetic and genomic medicine
10.64898/2026.02.20.26346743 medRxiv
Show abstract

Structural variants (SVs) are a major source of genomic diversity and disease susceptibility; however, populations from the Middle East and North Africa (MENA) region remain critically underrepresented in global reference databases. We provide the first detailed catalogue of structural variation in 61 individuals from diverse MENA countries, using publicly available ultra-long Oxford Nanopore sequencing. A scalable and dual-reference alignment-based method (GRCh38 and T2T-CHM13) was employed to comprehensively detect and characterise SVs in these samples. A robust multi-caller approach was applied to classify SVs as high-confidence true positives, requiring consensus from at least three callers. This approach identified 97,765 SVs using GRCh38 aXecting 11.6 Mb and 176,494 SVs against T2T-CHM13 aXecting 12.2 Mb, the latter providing improved alignment and variant detection. Significantly, up to 20% of the SVs identified in the MENA population were previously unreported, highlighting the regions uncharted genomic diversity. We identified population-specific structural variants within genes linked to diseases, pharmacogenetics, and immune responses, including SVs nearly fixed within MENA individuals overlapping with OMIM exons. Additionally, by integrating data from the 1K-ONT SV catalogue, alongside chimpanzee and archaic hominin genomes, shared ancient variants were identified, and MENA-specific SVs were highlighted. Evaluating the clinical utility of our resource for patients with MENA ancestry illustrated that it reduces the interpretation burden of SVs in these patients by 92%. These findings establish a foundational reference for structural variants relevant to MENA populations.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.