Back

Biostatistics

Oxford University Press (OUP)

All preprints, ranked by how well they match Biostatistics's content profile, based on 24 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Penalized longitudinal mixed models with latent group structure, with an application in neurodegenerative diseases

Hatami, F.; Perrakis, K.; Cooper-Knock, J.; Mukherjee, S.; Dondelinger, F.

2020-11-13 health informatics 10.1101/2020.11.10.20229302 medRxiv
Top 0.1%
53.3%
Show abstract

SO_SCPLOWUMMARYC_SCPLOWLarge-scale longitudinal data are often heterogeneous, spanning latent subgroups such as disease subtypes. In this paper, we present an approach called longitudinal joint cluster regression (LJCR) for penalized mixed modelling in the latent group setting. LJCR captures latent group structure via a mixture model that includes both the multivariate distribution of the covariates and a regression model for the response. The longitudinal dynamics of each individual are modeled using a random effect intercept and slope model. Inference is done via a profile likelihood approach that can handle high-dimensional covariates via ridge penalization. LJCR is motivated by questions in neurodegenerative disease research, where latent subgroups may reflect heterogeneity with respect to disease presentation, progression and diverse subject-specific factors. We study the performance of LJCR in the context of two longitudinal datasets: a simulation study and a study of amyotrophic lateral sclerosis (ALS). LJCR allows prediction of progression as well as identification of subgroups and subgroup-specific model parameters.

2
Semi-supervised Calibration of Risk with Noisy Event Times (SCORNET) Using Electronic Health Record Data

Ahuja, Y.; Liang, L.; Huang, S.; Cai, T.

2021-01-09 bioinformatics 10.1101/2021.01.08.425976 medRxiv
Top 0.1%
34.5%
Show abstract

Leveraging large-scale electronic health record (EHR) data to estimate survival curves for clinical events can enable more powerful risk estimation and comparative effectiveness research. However, use of EHR data is hindered by a lack of direct event times observations. Occurrence times of relevant diagnostic codes or target disease mentions in clinical notes are at best a good approximation of the true disease onset time. On the other hand, extracting precise information on the exact event time requires laborious manual chart review and is sometimes altogether infeasible due to a lack of detailed documentation. Current status labels - binary indicators of phenotype status during follow up - are significantly more efficient and feasible to compile, enabling more precise survival curve estimation given limited resources. Existing survival analysis methods using current status labels focus almost entirely on supervised estimation, and naive incorporation of unlabeled data into these methods may lead to biased results. In this paper we propose Semi-supervised Calibration of Risk with Noisy Event Times (SCORNET), which yields a consistent and efficient survival curve estimator by leveraging a small size of current status labels and a large size of imperfect surrogate features. In addition to providing theoretical justification of SCORNET, we demonstrate in both simulation and real-world EHR settings that SCORNET achieves efficiency akin to the parametric Weibull regression model, while also exhibiting non-parametric flexibility and relatively low empirical bias in a variety of generative settings.

3
Addressing selection bias in the UK Biobank neurological imaging cohort

Bradley, V. C.; Nichols, T. E.

2022-01-14 public and global health 10.1101/2022.01.13.22269266 medRxiv
Top 0.1%
28.6%
Show abstract

The UK Biobank is a national prospective study of half a million participants between the ages of 40 and 69 at the time of recruitment between 2006 and 2010, established to facilitate research on diseases of aging. The imaging cohort is a subset of UK Biobank participants who have agreed to undergo extensive additional imaging assessments. However, Fry et al. (2017) finds evidence of "healthy volunteer bias" in the UK Biobank - participants are less likely to smoke, be obese, consume alcohol daily than the target population of UK adults. Here we examine selection bias in the UK Biobank imaging cohort. We address two common misconceptions: first, that study size can compensate for bias in data collection, and second that selection bias does not affect estimates of associations, which are the primary interest of the UK Biobank. We introduce inverse probability weighting (IPW) as an approach commonly used in survey research that can be used to address selection bias in volunteer health studies like the UK Biobank. We discuss 6 such methods - five existing and one novel -, assess relative performance in simulation studies, and apply them to the UK Biobank imaging cohort. We find that our novel method, BART for predicting the probability of selection combined with raking, performs well relative to existing methods, and helps alleviate selection bias in the UK Biobank imaging cohort.

4
Large-Scale Sparse Regression for Multiple Responses with Applications to UK Biobank

Qian, J.; Tanigawa, Y.; Li, R.; Tibshirani, R.; Rivas, M. A.; Hastie, T.

2020-05-31 bioinformatics 10.1101/2020.05.30.125252 medRxiv
Top 0.1%
23.0%
Show abstract

In high-dimensional regression problems, often a relatively small subset of the features are relevant for predicting the outcome, and methods that impose sparsity on the solution are popular. When multiple correlated outcomes are available (multitask), reduced rank regression is an effective way to borrow strength and capture latent structures that underlie the data. Our proposal is motivated by the UK Biobank population-based cohort study, where we are faced with large-scale, ultrahigh-dimensional features, and have access to a large number of outcomes (phenotypes): lifestyle measures, biomarkers, and disease outcomes. We are hence led to fit sparse reduced-rank regression models, using computational strategies that allow us to scale to problems of this size. We use an iterative algorithm that alternates between solving the sparse regression problem and solving the reduced rank decomposition. For the sparse regression component, we propose a scalable iterative algorithm based on adaptive screening that leverages the sparsity assumption and enables us to focus on solving much smaller sub-problems. The full solution is reconstructed and tested via an optimality condition to make sure it is a valid solution for the original problem. We further extend the method to cope with practical issues such as the inclusion of confounding variables and imputation of missing values among the phenotypes. Experiments on both synthetic data and the UK Biobank data demonstrate the effectiveness of the method and the algorithm. We present multiSnpnet package, available at http://github.com/junyangq/multiSnpnet that works on top of PLINK2 files, which we anticipate to be a valuable tool for generating polygenic risk scores from human genetic studies.

5
A generalized synthetic control algorithm for sparse functional data

Shao, L.; Pohl, K. M.; Thompson, W. K.

2026-02-25 neuroscience 10.64898/2026.02.23.707582 medRxiv
Top 0.1%
22.7%
Show abstract

The Synthetic Control Method (SCM) and its interactive factor model generalizations (GSC) are powerful for estimating causal effects from panel data but are not easily applied when follow-up is irregular or sparse, common features of biomedical cohorts. We develop a Bayesian functional extension of GSC that treats each units outcome path as a smooth latent trajectory and accommodates unequally spaced measurements. Trajectories are approximated using Functional Principal Components Analysis (FPCA), providing a data-driven basis that captures dominant patterns with minimal shape assumptions while borrowing strength across individuals. Within this representation, we learn unit and time latent factors jointly with FPCA scores from the control data, construct counterfactual trajectories for treated units, and quantify uncertainty via the posterior. Identification relies on a latent-factor/weak-trend condition and overlap of controls and treated units in the functional score space. Simulation studies varying donor pool and treated unit size and sampling density show that the proposed approach (a.k.a GSC-FPCA) yields low bias when sampling is irregular or sparse, with well-calibrated interval coverage across a broad range of scenarios. We apply the method to longitudinal neuroimaging data from the National Consortium on Alcohol and Neurodevelopment in Adolescence - Adulthood (NCANDA-A) study to estimate the effect of adolescent binge drinking on subsequent brain volumes. Leveraging from 1 to 9 observed time points per participant, GSC-FPCA produces stable counterfactuals and detects a negative impact on gray-matter volumes with sustained high levels of binge drinking. Our results demonstrate that embedding GSC within a functional framework enables robust causal inference in biomedical applications characterized by irregularly-spaced visits, limited observations, and complex outcome dynamics.

6
Estimating Direct and Spillover Vaccine Effectiveness with Partial Interference under Test-Negative Design Sampling

Jiang, C.; Fang, F.; Talbot, D.; Schnitzer, M.

2025-02-25 infectious diseases 10.1101/2025.02.24.25322826 medRxiv
Top 0.1%
21.6%
Show abstract

The Test-Negative Design (TND), which involves recruiting care-seeking individuals who meet predefined clinical case criteria, offers valid statistical inference for Vaccine Effectiveness (VE) using data collected through passive surveillance, making it cost-efficient and timely. Infectious disease epidemiology often involves interference, where the treatment and/or outcome of one individual can affect the outcomes of others, rendering standard causal estimands ill-defined; ignoring such interference can bias VE evaluation and lead to ineffective vaccination policies. This article addresses the estimation of causal estimands for VE in the presence of partial interference using TND samples. Partial interference means that the vaccination of units within the same group/cluster may influence the outcomes of other members of the cluster. We define the population direct, spillover, total, and overall effects using the geometric risk ratio, which are identifiable under TND sampling. We investigate various stochastic policies for vaccine allocation in a counterfactual scenario, and identify policy-relevant VE causal estimands. We propose inverse-probability weighted (IPW) estimators for estimating the policy-relevant VE causal estimands with partial interference under the TND, and explore the statistical properties of these estimators.

7
Plasmodium vivax relapse, reinfection and recrudescence estimation using genetic data

Taylor, A. R.; Foo, Y. S.; White, M. T.

2022-11-24 infectious diseases 10.1101/2022.11.23.22282669 medRxiv
Top 0.1%
19.0%
Show abstract

Malaria is a major public health concern. Among the two most important causes of human malaria, Plasmodium vivax is the hardest to eliminate, owing largely to its ability to relapse (cause recurrent blood-stage infection following the activation of dormant liver-stage parasites called hypnozoites). Recurrent vivax malaria can also follow a failure to treat a preceding blood-stage infection (recrudescence) and, in endemic settings, a new infectious mosquito bite (reinfection). Understanding the cause of recurrent vivax malaria is critical for disease control efforts, e.g., to estimate the efficacy of a hypnozoiticidal drug, relapse needs to be separated from reinfection and recrudescence. In this report we describe the Pv3R statistical model designed to estimate using P. vivax genetic data the probability that a recurrent P. vivax blood-stage infection is a relapse, reinfection or recrudescence (the Pv 3Rs).

8
Stochastic epidemic models and their link with methods from survival analysis

Putter, H.; Goeman, J.; Wallinga, J.

2024-02-20 infectious diseases 10.1101/2024.02.18.24302991 medRxiv
Top 0.1%
18.6%
Show abstract

Compartmental models based on ordinary differential equations (ODEs) quantifying the interactions between susceptible, infectious, and recovered individuals within a population have played an important role in infectious disease modeling. The aim of the present paper is to explain the link between stochastic epidemic models based on the susceptible-infectious-recovered (SIR) model, and methods from survival analysis. We illustrate how standard software for survival analysis in the statistical language R can be used to estimate pivotal parameters in the stochastic SIR model in the very much idealized situation where the epidemic is completely observed. Extensions incorporating interventions, age structure and heterogeneity are explored and illustrated.

9
Generalization of the minimum covariance determinant algorithm for categorical and mixed data types

Beaton, D.; Sunderland, K. M.; ADNI, ; Levine, B.; Mandzia, J.; Masellis, M.; Swartz, R. H.; Troyer, A. K.; ONDRI, ; Binns, M. A.; Abdi, H.; Strother, S. C.

2020-03-31 bioinformatics 10.1101/333005 medRxiv
Top 0.1%
18.6%
Show abstract

The minimum covariance determinant (MCD) algorithm is one of the most common techniques to detect anomalous or outlying observations. The MCD algorithm depends on two features of multivariate data: the determinant of a matrix (i.e., geometric mean of the eigenvalues) and Mahalanobis distances (MD). While the MCD algorithm is commonly used, and has many extensions, the MCD is limited to analyses of quantitative data and more specifically data assumed to be continuous. One reason why the MCD does not extend to other data types such as categorical or ordinal data is because there is not a well-defined MD for data types other than continuous data. To address the lack of MCD-like techniques for categorical or mixed data we present a generalization of the MCD. To do so, we rely on a multivariate technique called correspondence analysis (CA). Through CA we can define MD via singular vectors and also compute the determinant from CAs eigenvalues. Here we define and illustrate a generalized MCD on categorical data and then show how our generalized MCD extends beyond categorical data to accommodate mixed data types (e.g., categorical, ordinal, and continuous). We illustrate this generalized MCD on data from two large scale projects: the Ontario Neurodegenerative Disease Research Initiative (ONDRI) and the Alzheimers Disease Neuroimaging Initiative (ADNI), with genetics (categorical), clinical instruments and surveys (categorical or ordinal), and neuroimaging (continuous) data. We also make R code and toy data available in order to illustrate our generalized MCD.

10
A Poisson reduced-rank regression model for association mapping in sequencing data

Fitzgerald, T.; Jones, A.; Engelhardt, B. E.

2022-06-01 bioinformatics 10.1101/2022.05.31.494236 medRxiv
Top 0.1%
18.6%
Show abstract

Single-cell RNA sequencing (scRNA-seq) technologies allow for the study of gene expression in individual cells. Often, it is of interest to understand how transcriptional activity is associated with cell-specific covariates, such as cell type, genotype, or measures of cell health. Traditional approaches for this type of association mapping assume independence between the outcome variables (or genes), and perform a separate regression for each. However, these methods are computationally costly and ignore the substantial correlation structure of gene expression. Furthermore, count-based scRNA-seq data pose challenges for traditional models based on Gaussian assumptions. We aim to resolve these issues by developing a reduced-rank regression model that identifies low-dimensional linear associations between a large number of cell-specific covariates and high-dimensional gene expression readouts. Our probabilistic model uses a Poisson likelihood in order to account for the unique structure of scRNA-seq counts. We demonstrate the performance of our model using simulations, and we apply our model to a scRNA-seq dataset, a spatial gene expression dataset, and a bulk RNA-seq dataset to show its behavior in three distinct analyses. We show that our statistical modeling approach, which is based on reduced-rank regression, captures associations between gene expression and cell- and sample-specific covariates by leveraging low-dimensional representations of transcriptional states.

11
Linear and partially linear models of behavioural trait variation using admixture regression

Connor, G.; Pesta, B. J.

2021-06-17 genomics 10.1101/2021.05.14.444173 medRxiv
Top 0.1%
18.5%
Show abstract

Admixture regression methodology exploits the natural experiment of random mating between individuals with different ancestral backgrounds to infer the environmental and genetic components to trait variation across racial and ethnic groups. This paper provides a statistical framework for admixture regression based on the linear polygenic index model and applies it to neuropsychological performance data from the Adolescent Brain Cognitive Development (ABCD) database. We develop and apply a new test of the differential impact of multi-racial identities on trait variation, an orthogonalization procedure for added explanatory variables, and a partially linear semiparametric functional form. We find a statistically significant genetic component to neuropsychological performance differences across racial identities, and find some possible evidence of nonlinearity in the link between admixture and neuropsychological performance scores in the ABCD data.

12
Hypothesis test of arbitrary parametric structure in a generalized additive model

Yang, Y.; Zhu, X.

2025-05-13 public and global health 10.1101/2025.05.12.25327450 medRxiv
Top 0.1%
18.3%
Show abstract

Before applying flexible nonparametric models such as a generalized additive model (GAM), it is natural to ask whether a simpler parametric form suffices. To address this question, we develop TAPS (Test for Arbitrary Parametric Structure), a framework that integrates estimation and hypothesis testing to evaluate whether a prespecified parametric form adequately captures a covariate effect in a GAM. TAPS accommodates diverse structures, including linearity, piecewise linearity with changepoints, discontinuities with jumps, among many others. It is implemented in the R package mgcv.taps built directly on mgcv, enabling seamless adoption, broad outcome support, and scalability to biobank-scale data. Using UK Biobank data, we analyze 38 continuous and 8 binary traits to investigate two scientific questions: does the effect of a polygenic risk score (PRS) vary with age beyond a linear interaction, and does retirement at age 65 modify this age-varying effect? We find that age-varying PRS effects are common and often strongly non-linear, and that retirement at 65 significantly modifies these effects for five traits after multiple-testing correction.

13
Four-Dimensional Sparse Bayesian Tensor Decomposition for Gene Expression Data

Gill, C. C.; Marchini, J.

2020-11-30 genetics 10.1101/2020.11.30.403907 medRxiv
Top 0.1%
18.2%
Show abstract

Disease etiology may be better understood through the study of gene expression in four dimensional (4D) experiments that consist of measurements on multiple individuals, genes, tissues and under multiple conditions or through time. We have developed a sparse Bayesian four dimensional tensor decomposition method aimed at uncovering latent components or gene networks that could be linked to genetic variation. We used a Variational Bayes algorithm to fit the model which provides fast and accurate analysis. In this brief note we illustrate the utility of the method using simulated datasets, and show that when 4D data is available our method shows improved performance in estimating the true structure in the dataset, when compared to using a 3D method on a single slice of the 4D dataset. We also compare the results of the 4D method to that of the 3D method on a suitable unfolding of the dataset, demonstrating that similar performance is observed in this case, while the 4D method accurately recovers the additional structure in the data. We provide software that implements the method in R.

14
Highly adaptive LASSO: Machine learning that provides valid nonparametric inference in realistic models

Butzin-Dozier, Z.; Qiu, S.; Hubbard, A. E.; Shi, J.; van der Laan, M.

2024-10-19 epidemiology 10.1101/2024.10.18.24315778 medRxiv
Top 0.1%
18.2%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWUnderstanding treatment effects on health-related outcomes using real-world data requires defining a causal parameter and imposing relevant identification assumptions to translate it into a statistical estimand. Semiparametric methods, like the targeted maximum likelihood estimator (TMLE), have been developed to construct asymptotically linear estimators of these parameters. To further establish the asymptotic efficiency of these estimators, two conditions must be met: 1) the relevant components of the data likelihood must fall within a Donsker class, and 2) the estimates of nuisance parameters must converge to their true values at a rate faster than n-1/4. The Highly Adaptive LASSO (HAL) satisfies these criteria by acting as an empirical risk minimizer within a class of cadlag functions with a bounded sectional variation norm, which is known to be Donsker. HAL achieves the desired rate of convergence, thereby guaranteeing the estimators asymptotic efficiency. The function class over which HAL minimizes its risk is flexible enough to capture realistic functions while maintaining the conditions for establishing efficiency. Additionally, HAL enables robust inference for non-pathwise differentiable parameters, such as the conditional average treatment effect (CATE) and causal dose-response curve, which are important in precision health. While these parameters are often considered in machine learning literature, these applications typically lack proper statistical inference. HAL addresses this gap by providing reliable statistical uncertainty quantification that is essential for informed decision-making in health research.

15
Multi-level Latent Variable Models for Coheritability Analysis in Electronic Health Records

Zhao, Y.; Tatonetti, N. P.; Wang, Y.

2025-11-06 public and global health 10.1101/2025.11.05.25339602 medRxiv
Top 0.1%
18.1%
Show abstract

Electronic health records (EHRs) linked with familial relationship data offer a unique opportunity to investigate the genetic architecture of complex phenotypes at scale. However, existing heritability and coheritability estimation methods often fail to account for the intricacies of familial correlation structures, heterogeneity across phenotype types, and computational scalability. We propose a robust and flexible statistical framework for jointly estimating heritability and genetic correlation among continuous and binary phenotypes in EHR-based family studies. Our approach builds on multi-level latent variable models to decompose phenotypic covariance into interpretable genetic and environmental components, incorporating both within- and between-family variations. We derive iteration algorithms based on generalized equation estimations (GEE) for estimation. Simulation studies under various parameter configurations demonstrate that our estimators are consistent and yield valid inference across a range of realistic settings. Applying our methods to real-world EHR data from a large, urban health system, we identify significant genetic correlations between mental health conditions and endocrine/metabolic phenotypes, supporting hypotheses of shared etiology. This work provides a scalable and rigorous framework for coheritability analysis in high-dimensional EHR data and facilitates the identification of shared genetic influences in complex disease networks.

16
Penalized generalized estimating equations for relative risk regression with applications to brain lesion data

Kindalova, P.; Veldsman, M.; Nichols, T. E.; Kosmidis, I.

2021-11-03 neuroscience 10.1101/2021.11.01.466751 medRxiv
Top 0.1%
17.1%
Show abstract

Motivated by a brain lesion application, we introduce penalized generalized estimating equations for relative risk regression for modelling correlated binary data. Brain lesions can have varying incidence across the brain and result in both rare and high incidence outcomes. As a result, odds ratios estimated from generalized estimating equations with logistic regression structures are not necessarily directly interpretable as relative risks. On the other hand, use of log-link regression structures with the binomial variance function may lead to estimation instabilities when event probabilities are close to 1. To circumvent such issues, we use generalized estimating equations with log-link regression structures with identity variance function and unknown dispersion parameter. Even in this setting, parameter estimates can be infinite, which we address by penalizing the generalized estimating functions with the gradient of the Jeffreys prior. Our findings from extensive simulation studies show significant improvement over the standard log-link generalized estimating equations by providing finite estimates and achieving convergence when boundary estimates occur. The real data application on UK Biobank brain lesion maps further reveals the instabilities of the standard log-link generalized estimating equations for a large-scale data set and demonstrates the clear interpretation of relative risk in clinical applications.

17
Incorporating LLM-Derived Information into Hypothesis Testing for Genomics Applications

Bryan, J. G.; Niu, H.; Li, D.

2025-05-07 bioinformatics 10.1101/2025.04.30.651464 medRxiv
Top 0.1%
15.5%
Show abstract

We propose strategies for incorporating the information in large language models (LLMs) into statistical hypothesis tests in genomics studies. Using gene embeddings derived from text inputs to OpenAIs GPT-3.5 model, we show that biological signals in a variety of genomics datasets reside near the principal subspace spanned by the embeddings. We then use a frequentist and Bayesian (FAB) framework to propose several hypothesis tests that are either optimal or approximately optimal with respect to prior information based on the gene embedding subspace. In four real-world genomics examples, the FAB tests guided by the LLM-derived information achieve more power than classical counterparts.

18
Estimation of total mediation effect for a binary trait in a case-control study for high-dimensional omics mediators

Kang, Z.; Chen, L.; Wei, P.; Xu, Z.; Li, C.; Yang, T.

2025-02-02 genomics 10.1101/2025.01.28.635396 medRxiv
Top 0.1%
15.1%
Show abstract

Mediation analysis helps uncover how exposures impact outcomes through intermediate variables. Traditional mean-based total mediation effect measures may suffer from the cancellation of opposite component-wise effects, and existing methods often lack the power to capture weak effects in high-dimensional mediators. Additionally, most high-dimensional mediation analysis methods have focused on continuous outcomes, with limited attention to binary outcomes, particularly in case-control studies. To fill this gap, we propose an R2 total mediation effect measure within the liability framework that offers a clear and intuitive causal interpretation, provides additional insights beyond the mean-based measures, and is invariant to disease prevalence. We develop a cross-fitted, modified Haseman-Elston regression-based estimation procedure tailored for mediation analysis in case-control studies, which can also be applied to cohort studies. Our estimator remains consistent in the presence of non-mediators and weak effects, as demonstrated in extensive simulations. Theoretical justification for consistency is provided under mild conditions and without requiring exact mediator selection. In a case-control substudy of the Womens Health Initiative involving 2150 individuals, we found that many metabolites were mediators with weak effects in the path from BMI to coronary heart disease, and we estimated that 89% (95% CI: 57%-100%) of the BMI-explained variation in underlying CHD liability is mediated by the measured metabolomics. The proposed estimation procedure is implemented in the R package "r2MedCausal", available on GitHub.

19
Mendelian Randomization with Instrumental Variable Synthesis (IVY)

Kuang, Z.; Cordova-Palomera, A.; Sala, F.; Wu, S.; Dunnmon, J.; Re, C.; Priest, J. R.

2019-09-09 bioinformatics 10.1101/657775 medRxiv
Top 0.1%
15.1%
Show abstract

Mendelian Randomization (MR) is an important causal inference method primarily used in biomedical research. This work applies contemporary techniques in machine learning to improve the robustness and power of traditional MR tools. By denoising and combining candidate genetic variants through techniques from unsupervised probabilistic graphical models, an influential latent instrumental variable is constructed for causal effect estimation. We present results on identifying relationships between biomarkers and the occurrence of coronary artery disease using individual-level real-world data from UK-BioBank via the proposed method. The approach, termed Instrumental Variable sYnthesis (IVY) is proposed as a complement to current methods, and is able to improve results based on allele scoring, particularly at moderate sample sizes.

20
Partitioning Fraction of Variance Explained into Strong Localized Effects and Weak Diffuse Effects

Nan, F.; Azriel, D.; Schwartzman, A.

2026-01-07 genetics 10.64898/2026.01.06.697735 medRxiv
Top 0.1%
15.0%
Show abstract

High-dimensional genetic data present substantial challenges for estimating the fraction of variance explained (FVE) by genome-wide single-nucleotide polymorphisms (SNPs). Standard approaches for SNP heritability estimation, such as GWAS heritability (GWASH) and linkage disequilibrium score (LDSC) regression, typically assume Gaussian distributions for SNP effect sizes. However, empirical evidence indicates that SNP effects are often heavy-tailed, with a small subset of variants exerting disproportionately large influence. Such settings violate the recently established bounded-kurtosis effect (BKE) condition, under which these FVE estimators are consistent. Consequently, widely used methods may yield severely biased estimates when strong effects are present. We introduce a decomposed FVE estimation framework that accommodates heavy-tailed and heterogeneous SNP effect distributions. The proposed approach partitions total heritability into contributions from strong and weak genetic effects, estimating the former using low-dimensional adjusted R2 and the latter using an extension of FVE estimation methodology that remains valid under BKE compliance. We further develop a test for detecting violations of the BKE condition and compare several high-dimensional screening procedures for identifying strong-effect SNPs when they are not known in advance. Simulation studies show that the proposed decomposition substantially improves estimation accuracy over existing approaches in the presence of heavy-tailed effects. Application to the Adolescent Brain Cognitive Development (ABCD) Study demonstrates the practical utility of the method, yielding more reliable heritability estimates for the PolyVoxel Score, a neuroimaging-based biomarker linked to iron accumulation. Our results highlight the importance of accommodating effect heterogeneity in large-scale genomic studies.