Back

The comprehensive detection of hemoglobinopathy variants via long-read sequencing

Xiang, J.; Peng, J.; Sun, X.; Jiang, C.; Zhao, H.; Guo, Y.; Xu, H.; Gu, S.; Ye, H.; You, L.; Huang, X.; Chen, S.; Zhu, B.; Peng, Z.

2024-12-06 genomics
10.1101/2024.12.03.626522 bioRxiv
Show abstract

BACKGROUNDThe genetic complexity of hemoglobin genes, characterized by high GC content and homologous sequences, poses significant challenges for detecting hemoglobin variants in clinical settings. METHODSA long-read indexed PCR method utilizing the novel CycloneSEQ nanopore sequencing platform was developed to detect all variant types, including single nucleotide variants (SNVs), deletions, structural variants (SVs) in HBA, HBB, HBD, and HBG genes. The method was validated using 507 clinical samples to assess its performance. RESULTSThe long-read indexed PCR system employed 13 primers targeting the hemoglobin gene clusters. This design enabled the detection of 37 types of HBA deletions, 5 SV (3 multicopies (, anti3.7, anti4.2) and 2 fusion allele (HK and anti-HK)), 37 HBB deletions, and all SNVs in the targeted regions. Validation across 507 samples (84 with HBA variants, 60 with HBB variants, 256 with both HBA and HBB variants, and 107 with no known variants) demonstrated 100.0% sensitivity and specificity. Additionally, the long-read sequencing enabled phasing of variants within hemoglobin genes, providing insights critical for clinical interpretation. CONCLUSIONSThe long-read indexed PCR method, combined with the CycloneSEQ nanopore sequencing platform, proved to be a robust and efficient solution for detecting hemoglobinopathy variants. The integration of indexed primers and barcoding enhances scalability, making this method ideal for large-scale population screening programs in the future.

Published in Human Genomics · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.