Forensic Science International: Genetics
○ Elsevier BV
All preprints, ranked by how well they match Forensic Science International: Genetics's content profile, based on 26 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Alsuwaidi, M. S.; Albastaki, A.; Almulla, H.; Omar, A. K.; Almarri, M. A.
Show abstract
Short tandem repeat (STR) profiling is the cornerstone of forensic DNA analysis, traditionally performed via capillary electrophoresis. Recently, next generation sequencing has gained prominence due to its increased discriminatory power and enhanced performance with degraded samples. Nanopore sequencing offers a portable and cost-effective alternative, however historically high error rates have precluded its forensic adoption. Here, we evaluate R10.4.1 flow cell chemistry and multiple basecalling tiers (HAC, SUP, HYP) across several iterations (v4.2, v5.0, v5.2, v6.0) to assess their impact on genotyping accuracy. Analyzing 45 STR loci (22 autosomal and 23 Y-STRs) across single-source controls, we introduce a parallelized, user-friendly pipeline designed to transform raw POD5 files into STR profiles. Our results demonstrate a progressive improvement in genotyping accuracy with each basecalling iteration, with the latest models achieving 99.0% autosomal and 100% Y-STR concordance. Furthermore, we find that filtering on raw-read quality scores significantly improves genotyping by reducing background noise and generating cleaner profiles. Notably, the HYPv5.0 Q20 filter drove an average 53.2% reduction in misaligned reads across all loci in comparison to earlier basecalling models. Our study demonstrates that continual bioinformatic improvements in basecalling models, coupled with R10.4.1 chemistry, can provide accurate STR profiles in single-source samples, warranting larger validation studies with more diverse samples to further evaluate performance.
Berger, J.; Krawczak, M.; Zandstra, D.; Kayser, M.; Ralf, A.; Scheurer, E.; Caliebe, A.; Schulz, I.
Show abstract
The formal assessment of a genetic match between a suspect and some biological trace material is one of the key tasks of forensic genetics, particularly in cases of sexual offence. The analysis of Y-chromosomal short tandem repeats (Y-STRs) has proven especially useful in this context. For a long time, however, calculating the probability of a perfect Y-STR profile match under the defense hypothesis that the suspect was not the trace donor posed a great challenge. This was due to the inherent uncertainty about the population of alternative donors, the so-called suspect population. We recently proposed to resolve this controversy by systematically favoring the suspect and considering his close male relatives as the suspect population. However, since the mathematical framework developed for this purpose was simulation-based, its practical application turned out increasingly difficult with increasing pedigree size. Here, we present an adaptation of the so-called Elston-Stewart algorithm, originally developed for the linkage analysis of human genetic diseases, to allow calculation of exact match probabilities in a time that scales linearly with pedigree size. The adapted algorithm was implemented in a publicly available software tool, and its correctness was verified by the comparison of its output with the correct, analytical results obtained for selected example pedigrees. The new implementation mostly outperforms the simulation-based solution, albeit with the important exception of Y-STRs present in multiple copies. Given the increasingly prominent role of such multicopy markers in forensic genetics, the complementary use of both approaches appears the most sensible strategy for the time being.
Pascali, V. L.
Show abstract
Single nucleotide polymorphisms (SNPs) are useful forensic markers. When a SNPs-based forensic protocol targets a body fluid stain, it returns elementary evidence regardless of the number of individuals that might have contributed to the stain deposition. Therefore, drawing inference from a mixed stain with SNPs is different than drawing it while using multinomial polymorphisms. We here revisit this subject, with a view to contribute to a fresher insight into it. First, we manage to model conditional semi-continuous likelihoods in terms of matrices of genotype permutations vs number of contributors (NTZsc). Secondly, we redefine some algebraic formulas to approach the semi-continuous calculation. To address allelic dropouts, we introduce a peak height ratio index ( h, or: the minor read divided by the major read at any NGS-based typing result) into the semi-continuous formulas, for they to act as an acceptable proxy of the split drop (Haned et al, 2012) model of calculation. Secondly, we introduce a new, empirical method to deduct the expected quantitative ratio at which the contributors of a mixture have originally mixed and the observed ratio generated by each genotype combination at each locus. Compliance between observed and expected quantity ratios is measured in terms of (1-{chi}2) values at each state of a locus deconvolution. These probability values are multiplied, along with the h index, to the relevant population probabilities to weigh the overall plausibility of each combination according to the quantitative perspective. We compare calculation performances of our empirical procedure (NITZq) with those of the EUROFORMIX software ver. 3.0.3. NITZq generates LR values a few orders of magnitude lower than EUROFORMIX when true contributors are used as POIs, but much lower LR values when false contributors are used as POIs. NITZ calculation routines may be useful, especially in combination with mass genomics typing protocols.
Gettings, K. B.; Tillmar, A.; Marshall, C.; Sturk-Andreaggi, K.
Show abstract
In mass disaster events, forensic DNA laboratories may be called upon to quickly pivot their operations toward identifying bodies and reuniting remains with family members. Ideally, laboratories have considered this possibility in advance and have a plan in place. Compared with traditional short tandem repeat (STR) typing, single nucleotide polymorphisms (SNPs) may be better suited to these disaster victim identification (DVI) scenarios due to their small genomic target size, resulting in an improved success rate in degraded DNA samples. As the landscape of technology has shifted toward DNA sequencing, many forensic laboratories now have benchtop instruments available for massively parallel sequencing (MPS), facilitating this operational pivot from routine forensic STR casework to DVI SNP typing. Herein, we review the commercially available SNP sequencing assays amenable to DVI, we use data simulations to explore the potential for kinship prediction from SNP panels of varying size, and we give an example DVI scenario as context for presenting the matrix of considerations: kinship predictive potential, cost, and throughput of current SNP assay options. This information is intended to assist laboratories in choosing a SNP system for disaster preparedness. Highlights3 to 5 bullet points (maximum 100 characters per bullet point, including spaces). Each bullet point should be a full sentence and should outline the key contributions of your manuscript and how it impacts forensic science. O_LISingle nucleotide polymorphisms (SNPs) are useful in disaster victim identification (DVI). C_LIO_LISNP panels amenable to human identification and extended kinship are described. C_LIO_LISimulations demonstrate the potential for kinship prediction from SNP panels of varying size. C_LIO_LIKinship predictive potential, cost, and throughput are presented for an example DVI scenario. C_LIO_LIInformation is intended to assist laboratories in choosing a SNP system for disaster preparedness. C_LI
Ismach, H.; Greenbaum, G.; Kennett, D.; Carmi, S.
Show abstract
Forensic investigative genetic genealogy (FIGG) is a revolutionary method in forensic genetics, whereby genetic relatives of an unknown target person are detected in direct-to-consumer genomic databases; the family trees of these relatives are then reconstructed to suggest candidates for the target person. Despite recent successes, the full potential of the technology has not yet been systematically evaluated at the scale of an entire country. To estimate the proportion of FIGG cases where a genetic relative is detected in the database (the "match rate"), we used the Israeli population register, covering all past and present citizens. After extensive quality control, the register included 12.4 million individuals, among them 10.4 million alive. We simulated genomic databases by randomly selecting subsets of predefined sizes of the live adult population. With a database covering 1% of the population, we estimate that FIGG would detect at least one relative of fifth degree (e.g., a second cousin) or closer for about 25% of the population, and at least two relatives for 9% of the population. A database covering 5% of the population would find at least one relative of third degree (e.g., a first cousin) or closer for half the population. More distant relatives are rarely identified in the register, likely due to its limited time depth. Our results provide the first country-wide direct estimates for the utility of FIGG in generating investigative leads.
Egeland, T.; Marsico, F.
Show abstract
Kinship cases, ranging from standard paternity tests to complex disaster victim identifications, are typically evaluated using likelihood ratios (LR) based on forensic genetic markers. However, in some contexts, genetic information alone is not enough to reach conclusive results. This is common when establishing distant familial connections using large DNA-databases, or even in simple cases such as determining which individual is the parent and which is the child in a relationship pair. Although forensic practitioners frequently incorporate additional evidence (SE), such as age, biological sex, or phenotypic traits, in these cases, this integration typically occurs informally, without rigorous probability estimation, compromising procedural transparency and reliability. Here, we present a comprehensive methodological framework that formally synthesizes forensic DNA evidence (FDE) with SE through Markov chain models and customized transition matrices designed for various biological traits. This approach generates combined likelihood assessments expressed as LRs or posterior probabilities. Validation through simulated and real-world case studies demonstrates that systematic incorporation of SE improves resolution accuracy in kinship determinations. To facilitate adoption, we have implemented this methodology in mispitools, an open-source R package.
Zhu, B.; Li, D.; Han, G.; Yao, X.; Gu, H.; Liu, T.; Liu, L.; Dai, J.; Liu, I. Z.; Liang, Y.; Zheng, J.; Sun, Z.; Lin, H.; Wang, W.; Liu, N.; Yu, H.; Shi, M.; Shen, G.; Qu, L.
Show abstract
Estimation of chronological age is particularly informative in forensic contexts. Assessment of DNA methylation status allows for the prediction of age, though the accuracy and ease of manipulation may vary across different models. In this study, we started with a carefully designed discovery cohort recruiting more elderly subjects than other age categories, to diminish the effect of epigenetic drifting. We analyzed DNA methylation from a single genomic region of ELOV2, which was sufficient to construct an age-prediction model comprising 15 CpG sites. This model is further validated by an independent cohort as well as a multi-center test using trace dried bloodstains. The nature of our analytical pipeline, when combined the assessment of a single genomic locus with high-throughput sequencing, can easily be scaled up with low cost. Taken together, we propose a new age-prediction model featuring accuracy, ease of manipulation, high-throughput, and low cost. This model can be readily applied in both classic and newly emergent forensic contexts that require age estimation.
Marsico, F.; Egeland, T.
Show abstract
Recent years have seen significant advances in DNA phenotyping, which predicts the physical traits of an unknown person, such as hair, eyes, and skin color, using DNA data. This technique is increasingly used in forensic investigations to identify missing persons, disaster victims, and suspects of crimes. A key contribution of DNA phenotyping is that it allows researchers to search through lists of individuals with similar characteristics, often gathered from testimonies, photographs, and social media data. However, despite their growing relevance, current methods lack comprehensive mathematical models to calculate likelihood ratios that accurately assess the statistical weight of evidence. Our work bridges this gap by developing new likelihood ratio models, validated through computational simulations. In addition, we demonstrate the ability of these models to improve forensic investigations in real-world scenarios. Furthermore, we introduce the R package forensicolors, freely available on CRAN, to facilitate the application of the methodologies developed.
Gill, P.; Bleka, O.
Show abstract
The interpretation of findings of low-template DNA given activity-level propositions requires robust statistical models capable of accommodating substantial inter-laboratory and case-specific variability. This paper presents the practical implementation of HaloGen, an open-source hierarchical Bayesian framework for calculating activity-level likelihood ratios (LRs) from DNA quantity data. We compare three modelling approaches derived from the framework: a Group model, which combines data across laboratories, a hierarchically informed Lab-Bayes model, and a standalone, laboratory specific Lab-Vague model. Through a series of simulation studies, we demonstrate that evidential strength is highly sensitive not only to DNA quantity but also to case context, particularly the assumed number of offenders (NS). We further show that inter-laboratory differences in DNA recovery and dropout can lead to materially different LRs, making unvalidated use of pooled or external data potentially misleading. To address practical implementation, we propose a minimum-effort validation pathway for laboratories wanting to report findings given activity level propositions. Our results indicate that a small number of direct/secondary transfer experiments (n {approx} 6- 12) are sufficient to obtain conservative LRs compared with a generic population model. Finally, these results clarify how contextual assumptions enter mathematically into activity-level inference, demonstrating that confirmation bias can arise naturally from unexamined modelling choices and underscoring the importance of transparent, explicit specification of propositions and parameters.
Ridings, R.; Gabriel, A.; Elliott, C. I.; Shafer, A.
Show abstract
DNA quantification technology has increased in accuracy and sensitivity, now allowing for detection and profiling of trace DNA. Secondary DNA transfer occurs when DNA is deposited via an intermediary source (e.g. clothing, tools, utensils). Multiple courtrooms have now seen secondary transfer introduced as an explanation for DNA being present at a crime scene, but sparse experimental studies mean expert opinions are often limited. Here, we used bovine blood and indigo denim substrates to quantify the amount of secondary DNA transfer and quality of STRs under three different physical contact scenarios: passive, pressure, and friction. We showed that the DNA transfer was highest under a friction scenario, followed by pressure and passive treatments. The STR profiles showed a similar, albeit less pronounced trend, with correctly scored alleles and genotype completeness being highest under a friction scenario, followed by pressure and passive. DNA on the primary substrate showed a decrease in concentration and genotype completeness both immediately and at 24 hours, suggestive of a loss of DNA during the primary transfer. The majority of secondary transfer samples amplified less than 50% of STR loci regardless of contact type. This study showed that while DNA transfer is common between denim, this is not manifested in full STR profiles. We discuss the possible technical solutions to partial profiles from trace DNA, and more broadly the ubiquity of secondary DNA transfer.
Gill, P.; Bleka, O.
Show abstract
The interpretation of trace DNA evidence at activity level requires explicit modelling of transfer, persistence, and failure to detect a person of interest. We present the theoretical foundations of HaloGen, an open-source hierarchical Bayesian framework for evaluating biological results under competing activity-level propositions, such as direct versus secondary transfer. HaloGen accounts for dropout, multiple contributors, and multiple stains. Evidence is evaluated using an exhaustive-propositions likelihood ratio frame-work that combines information across contributors and stains, while fully accounting for uncertainty in transfer and detection. Observed DNA quantities and non-detects are handled consistently within a single probabilistic model, avoiding reliance on fixed parameter estimates. The framework yields intuitive and robust behaviour: strong support for direct transfer when DNA quantities are informative, and appropriately neutral or defence-leaning likelihood ratios in low-information or non-detect scenarios. An empirically constrained fail-rate parameter prevents spurious inflation of likelihood ratios when offender detection is unlikely, providing stability across laboratories and experimental conditions. This paper establishes the theoretical basis of HaloGen; a companion paper addresses validation and applied casework examples.
Flores, M.; Pellegrini, M.
Show abstract
Chronological age estimation can provide supporting information in forensic casework when traditional identification methods are limited. DNA methylation, a stable epigenetic mark, has emerged as a promising tool for predicting chronological age from trace samples. However, many existing age estimation models rely on linear regression approaches, which often yield biased prediction errors across the age distribution (i.e. model residuals show a significant age dependence). In this study, we compared three approaches for age estimation modeling: multivariable linear regression, random forest regression and maximum likelihood estimation. While the first two approaches are well established, for the third one we constructed and validated a DNA methylation-based LOESS regression maximum likelihood model for age estimation utilizing forensic-relevant CpG markers. In all cases, model performance was evaluated through Leave-One-Out Cross-Validation (LOOCV). We utilized three independent publicly accessible methylation datasets collected using droplet digital PCR (ddPCR) to evaluate the most effective method for accuracy and bias in age estimation. Notably, when we compare the results of the maximum likelihood approach to the other approaches, multivariable linear regression and random forest regression, we find less bias in the age associated residuals compared to the other methods. These findings highlight the utility of non-linear modeling techniques in reducing the biases of epigenetic age estimation for forensic applications.
Vol, E.; Waldman, S.; Lomes, A.; Brielle, E. S.; Appel, N.; Dolin, B.; Asif, S.; Nagar, Y.; Marco, E.; Bergman, N.; Khaner, O.; Raviv, D.; Oliel, J.; Lewis, R. Y.; Carmi, S.
Show abstract
Genome-wide technologies can generate investigative leads in cold cases by determining the genetic ancestry of the forensic sample. Increasingly, DNA extraction and whole-genome sequencing or genotyping are being used to analyze early or middle-20th century skeletal remains. Here, we present the first case, to our knowledge, of whole-genome sequencing of a middle-20th-century bone sample from the Middle East. A femur discovered in a cave in Central Israel was proposed to belong to a person of Ashkenazi Jewish ancestry who was missing since 1948. Following DNA extraction and single-stranded library preparation, whole-genome sequencing generated nearly 500 million reads. However, only 0.5% of the reads mapped to the human genome, providing depth of coverage of 0.07x. After quality control and male sex inference, ancestry assignment was performed using principal components and ADMIXTURE analyses. The results suggested that the genome definitively belonged to a person of Arab ancestry, refuting the hypothesis of an Ashkenazi Jewish origin.
Poggiali, B.; Aagreen, C. I. V.; Meyer, O. L.; Jepsen, A. H.; Korneliussen, T. S.; Kampmann, M.-L.; Borsting, C.; Andersen, J. D.
Show abstract
Shotgun sequencing (SGS) enables simultaneous interrogation of a broad range of loci across the human genome, even from low-template and highly degraded DNA samples. While human identification traditionally relies on short tandem repeats (STRs) due to their high polymorphism, standard forensic STRs (100-450 bp) are poorly suited for the short read ([~]150 bp) constraint of SGS. The purpose of this study was to evaluate the analysis limitations of standard forensic STRs in SGS data and to identify a novel panel of STRs optimised for short-read genomic data. First, we benchmarked four STR genotyping software tools (STRait Razor, GangSTR, STRinNGS, and HipSTR) by analysing 53 standard forensic STRs in SGS data. HipSTR showed the best performance but achieved only a call rate of 64.5% and an accuracy of 83.8%, and its performance was strongly affected by STR allele length and read depth. To overcome these constraints, we screened the population-wide 1000 Genomes Project dataset and identified a panel of 265 autosomal ultra-short (< 50 bp) STRs with an effective number of alleles (Ae) ranging from 3.0 to 7.5. As few as seven of these loci were sufficient to achieve a Mean Match Probability (MMP) below 1 x 10-6. To validate these findings, we developed a custom PCR-based amplicon sequencing panel targeting 97 of the most polymorphic ultra-short STRs and evaluated these in 41 blood samples from Danish individuals. The polymorphic nature of the selected loci was confirmed (Aeranged from 2.4 to 7.2). Our results furthermore demonstrated high concordance between the amplicon panel and SGS-derived genotypes, which substantiates that these ultra-short STRs provide a robust and highly polymorphic alternative for human identification in SGS data. Author summaryShotgun sequencing (SGS) methods are increasingly being adopted in fields such as forensic genetics. SGS yields large amounts of genetic information by reading short fragments across the entire genome, enabling a wide range of analyses that may be exploited as leads in a police investigation. Human identification has traditionally been based on STR loci with a PCR amplicon length of 100-450 base pairs. However, these loci are often longer than the reads generated by SGS data, which makes them difficult to analyse in a reliable way. In this study, we evaluated four software tools designed to genotype STRs and confirmed the limited ability to genotype traditional forensic STRs in SGS data. To address this limitation, we identified a new set of highly polymorphic ultra-short STRs (less than 50 base pairs in length) that enable robust human identification using SGS data. Despite their shorter length, these loci retain the multi-allelic nature inherent to traditional STRs. This ensures a low random match probability that is comparable with the standard forensic STR panels. The ultra-short STRs may be genotyped from highly degraded DNA and may provide the possibility for complex mixture analysis and multi-donor deconvolution, which makes the STRs uniquely suited for forensic casework.
Swayambhu, M.; Gysi, M.; Haas, C.; Schuh, L.; Walser, L.; Javanmard, F.; Flury, T.; Ahannach, S.; Lebeer, S.; Hanssen, E. N.; Snipen, L.; Bokulich, N.; Kuemmerli, R.; Arora, N.
Show abstract
BackgroundRecent advances in next-generation sequencing have opened up new possibilities for utilizing the human microbiome in various fields, including forensics. Researchers have capitalized on the site-specific microbial communities found in different parts of the body to identify body fluids from biological evidence. Despite promising results, microbiome-based methods have not yet been fully integrated into forensic practice due to the lack of standardized protocols and systematic testing of methods on forensically relevant samples. Our study addresses critical decisions in establishing these protocols, focusing on bioinformatics choices and the use of machine learning to present microbiome results in court for forensically relevant and challenging samples. ResultsWe propose using Operational Taxonomic Units (OTUs) for read data processing and creating heterogeneous training datasets for training a random forest classifier. Our classifier incorporates six forensically relevant classes: saliva, semen, hand skin, penile skin, urine, and vaginal/menstrual fluid. Across these classes, our classifier achieved a high weighted average F1 score of 0.89. Systematic testing on mixed-source samples and underwear revealed reliable detection of at least one component of the mixture and the identification of vaginal fluid from underwear substrates. Additionally, when investigating the sexually shared microbiome (sexome) of heterosexual couples, our classifier shows promising results for the inference of sexual activity. ConclusionIn our study, we recommend the use of a novel random forest classifier trained on a heterogenous dataset for obtaining predictions from samples mimicking forensic evidence. We also highlight the potential of the sexome for assessing the nature of sexual activities in forensic investigations, while delineating areas that warrant further research. Furthermore, we underscore key considerations when presenting machine learning results for classifying mixed-source samples.
Frontanilla, T. S.; Valle Silva, G.; Ayala, J.; Mendes, C. T.
Show abstract
Accurate STR genotyping from next-generation sequencing (NGS) data has been challenging. Haplotype inference and phasing for STRs (HipSTR) was specifically developed to deal with genotyping errors and obtain reliable STR genotypes from whole-genome sequencing datasets. The objective of this investigation was to perform a comprehensive genotyping analysis of a set of STRs of broad forensic interest from the 1000 Genomes populations and release a reliable open-access STR database to the forensic genetics community. A set of 22 STR markers were analyzed using the CRAM files of the 1000 Genomes Project Phase 3 high-coverage (30x) dataset generated by the New York Genome Center (NYGC). HipSTR was used to call genotypes from 2,504 samples from 26 populations organized into five groups: African, East Asian, European, South Asian, and admixed American. The D21S11 marker could not be detected in the present study. Moreover, the Hardy-Weinberg equilibrium analysis, coupled with a comprehensive analysis of allele frequencies, revealed that HipSTR could not identify longer Penta E (and Penta D at a lesser extent) alleles. This issue is probably due to the limited length of sequencing reads available for genotype calling, resulting in heterozygote deficiency. Notwithstanding that, AMOVA, a clustering analysis using STRUCTURE, and a Principal Coordinates Analysis revealed a clear-cut separation between the four major ancestries sampled by the 1000 Genomes Consortium (AFR, EUR, EAS, SAS). Meanwhile, the AMOVA results corroborated previous reports that most of the variance is (97.12%) observed within populations. This set of analyses revealed that except for larger Penta D and Penta E alleles, allele frequencies and genotypes defined by HipSTR from the 1000 Genomes Project phase 3 data and offered as an open-access database are consistent and highly reliable.
Aneli, S.; Nicolini, V.; Vincenti, G.; Montinaro, F.; Sasso, S.; Saupe, T.; Kabral, H.; Solnik, A.; Tambets, K.; Guglielmino, R.; Fabbri, P. F.; Pagani, L.
Show abstract
BackgroundRoca Vecchia, an iconic Bronze Age stronghold in Apulia, Southern Italy, was completely destroyed during a siege between the end of the 15th century BCE and the beginning of the 14th century BCE. During the siege, seven of the local people hid within the stronghold walls. Two others, who could have been as well attackers as defenders, were found under the ruins of the main gate. The material culture found at Roca Vecchia and associated with the period of the siege includes Minoan-type pottery produced from local clay, imported Aegean pottery and an Aegean-type dagger, pointing to an established relationship between the site and the Minoan civilization. Therefore, the site offers an unprecedented opportunity to characterise the genetic components of the population inhabiting an indigenous settlement with increasing contacts with the Aegean world, and to shed light on the demic or cultural modes of the Minoan presence in the central Mediterranean. ResultsWith our work, we sampled six out of nine available unburied Middle Bronze Age individuals, contemporary with the siege and destruction of the site, and obtained genome-wide information for two individuals. When compared with available Minoan, Apulian and broadly Mediterranean genomes, the individuals showed a characteristic Bronze/Iron Age Italian peninsula genetic signature, with limited contribution from Minoans. ConclusionsWe conclude that the local population of Roca Vecchia, at the moment of the siege, was predominantly autochthonous, with a minoritarian Minoan component. A Minoan genetic signal is indeed likely present in one out of two analysed individuals who were certainly part of the dwellers of Roca Vecchia. This confirms previous hypotheses supposing that a nucleus of "foreigners" coming from the Minoan world was living in the site and mixed with locals. Archaeological data suggest that the Roca Vecchia Aegean population component probably increased in the following centuries.
Gosch, A.; Courts, C.
Show abstract
Interest in forensic RNA analysis has increased over the last years. RNA molecules present in forensic samples can accurately be quantified via quantitative PCR (qPCR), however, due to the limited number of markers that can be assayed simultaneously per reaction, qPCR is less suitable for applications requiring gene expression quantification of large marker sets. Few years ago, massively parallel targeted RNA-sequencing (targRNAseq) allowing to simultaneously and accurately quantify several hundreds of markers has been added to the forensic genetic tool set. However, typical targRNAseq protocols include a multiplex-PCR-step to amplify selected targets which potentially introduces bias and limits accurate gene expression quantification. Unique Molecular Indices (UMIs) have been invented to overcome this limitation and have been implemented in protocols from some vendors. In this study, we compared two targeted RNAseq protocols assaying expression of a set of 121 forensically relevant mRNA biomarkers: The Ion Ampliseq targeted RNA sequencing panel (Thermo Fisher Scientific), which employs a multiplex-PCR without the use of UMIs, and the QIAseq targeted RNA panel (QIAGEN), which uses UMIs prior to multiplex amplification. Both protocols were tested on replicated samples and dilution series and compared with respect to sensitivity and accuracy of gene expression quantification. The UMI-based protocol exhibited decreased sensitivity in comparison to the non-UMI-based alternative, however, making use of UMI technology greatly improved gene expression quantification accuracy. We thus recommend the use of UMI-based protocols for targeted RNA sequencing for applications requiring accurate gene expression quantification.
Wang, X.; King, J.; Meng, H.; Coble, M. D.; Woerner, A. E.
Show abstract
Variant calling is a ubiquitous genomic technique that underpins many scientific disciplines. From a computational perspective, variant calling is a form of logical compression; neglecting large variation, a persons genome can be losslessly described as a set of differences (SNP and small InDel alleles) relative to the reference sequence. Another common genomic technique is haplotype phasing, wherein alleles are partitioned into their paternal and maternal components (as haplotypes). Some classes of alleles are more difficult to describe than others, e.g., short tandem repeats (STRs). STRs serve as a critical marker for many genetic assays. However, STRs tend not to be explicitly reported in most genomic workflows. Here, we present StrPhaser, a novel algorithm that leverages phased variant calling datasets in the VCF file format to construct STR alleles. We evaluated StrPhaser on [~]10,000 STR alleles from 284 human genomes, achieving an average allele accuracy of 91%. In addition, StrPhaser better recovers longer STR alleles than competing approaches; in principle, STR alleles that are longer than the maximum read length can be characterized. This capability, combined with its user-friendly interface, speed, and generation of both STR genotypes and visualizations, makes StrPhaser a valuable tool for a wide range of genomic studies. AvailabilityThe StrPhaser is publicly available at https://github.com/XuewenWangUGA/StrPhaser.
Bright, J.-A.; Lee, S.-I.; BUCKLETON, J.; Taylor, D. A.
Show abstract
In previously reported work a method for applying a lower bound to the variation induced by the Monte Carlo effect was trialled. This is implemented in the widely used probabilistic genotyping system, STRmix. The approach did not give the desired 99% coverage. However, the method for assigning the lower bound to the MCMC variability is only one of a number of layers of conservativism applied in a typical application. We tested all but one of these sources of variability collectively and term the result the near global coverage. The near global coverage for all tested samples was greater than 99.5% for inclusionary average LRs of known donors. This suggests that when included in the probability interval method the other layers of conservativism are more than adequate to compensate for the intermittent underperformance of the MCMC variability component. Running for extended MCMC accepts was also shown to result in improved precision.