Statistics in Medicine
○ Wiley
All preprints, ranked by how well they match Statistics in Medicine's content profile, based on 40 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Chen, Y.; Gui, T.; Huang, Z.; Quach, N.; Tu, S.; Liu, J.; Garrett, T. J.; Starkweather, A. R.; Lyon, D. E.; Shepherd, B. E.; Tu, X. M.; Lin, T.
Show abstract
SO_SCPLOWUMMARYC_SCPLOWChemotherapy in breast cancer (BC) can substantially affect mental wellness. Advances in metabolomics enable comprehensive profiling of metabolic changes over time during and after treatment, offering insights into biological mechanisms linking chemotherapy to mental health outcomes. To study the association between metabolite profiles and mental wellness, correlation-based analyses are particularly useful. Spearmans rho is a widely used correlation measure and popular alternative to Pearsons correlation, since it also applies to non-linear association between variables. However, existing methods are not designed for longitudinal data and do not allow for covariate adjustments. In this paper, we propose a novel regression-based framework grounded in a class of semiparametric models, the functional response models, to extend this popular correlation measure to longitudinal settings with missing data under the missing at random assumption. This framework facilitates inferences about temporal changes in correlations over time and association of explanatory variables for such changes. We use simulation studies to evaluate performance of the approach with moderate sample sizes. We apply the approach to a one-year longitudinal substudy of the EPIGEN study to examine the longitudinal association between metabolite profiles and mental wellness in BC patients undergoing chemotherapy. The identified metabolites may serve as candidates for future in-depth bioinformatics analyses and translational investigations.
Turkia, J.; Schwab, U.; Hautamäki, V.
Show abstract
Maintaining proper nutrition is crucial for preserving health and preventing disease. However, what constitutes proper nutrition may vary among individuals; evidence indicates that the effects of diet and even single nutrients can differ considerably because of personal characteristics. This personal variability can be observed through blood markers, such as concentrations of plasma cholesterol and insulin, and captured using a hierarchical multivariate model. We leverage this variability and propose a conditional two-component Bayesian mixture model for generating personalized diet recommendations. The model uses the Nordic Nutrition Recommendations 2023 as a prior for healthy intake and infers individualized recommendations as posterior distributions. The first component identifies dietary options predicted to produce healthy levels across all considered blood markers, while the second selects, among these valid options, the diet closest to predefined personal preferences. The preference component is configurable and, in this study, was used to minimize dietary adjustments to support recommendation adherence while providing well-defined targets for nutrients less critical to concentration regulation. The method was evaluated using nutritional data from two studies: one in prediabetic individuals and one in patients with kidney dysfunction. Numerical simulations showed that the individualized diets could restore or approach normal plasma concentrations when the estimated personal nutrient effects indicated biological feasibility. As the results align with current nutritional literature, the Bayesian approach offers a principled way to leverage observational nutrition data. However, future clinical studies are needed to validate the results and modeling approach before these can be translated into evidence-based personalized nutritional counseling.
Stijven, F.; Mallinckrodt, C. H.; Molenberghs, G.; Alonso, A.; Dickson, S.; Hendrix, S.
Show abstract
In progressive diseases, like Alzheimers disease, treatments that slow progression should start early in the disease course to longer maintain higher levels of functioning. In corresponding clinical trials, the treatment effect is usually expressed in terms of mean differences on a clinical scale. Early in the disease course, however, treatment effects expressed on a clinical scale are often small but may nonetheless correspond to an important slowing of disease progression. This complicates the appreciation of the relevance of observed treatment effects. For example, it may be difficult to determine whether a 2-point improvement on a clinical scale is relevant for clinical practice. In this paper, we propose the meta Time-Component Tests (meta TCT). This new approach leads to estimators of treatment effects on the time scale, in terms of time saved or percentage slowing of progression, that are easy to interpret. This approach is based on estimates obtained from an arbitrary model for longitudinal data and is, therefore, very flexible. Asymptotic properties of the Meta TCT estimators are derived and evaluated in an extensive simulation study. Meta TCT is then applied to a phase 2/3 clinical trial for Alzheimers disease, which was first analyzed with a mixed model. In this trial, meta TCT leads to important additional insights into the treatment effect. We believe that meta TCT will facilitate the estimation of interpretable treatment effects in clinical trials for progressive diseases, and that this, in turn, will fine-tune the evaluation of the clinical relevance of new treatments.
Xu, Z.; Li, C.; Chi, S.; Yang, T.; Wei, P.
Show abstract
Mediation analysis is a useful tool in investigating how molecular phenotypes such as gene expression mediate the effect of exposure on health outcomes. However, commonly used mean-based total mediation effect measures may suffer from cancellation of component-wise mediation effects in opposite directions in the presence of high-dimensional omics mediators. To overcome this limitation, we recently proposed a variance-based R-squared total mediation effect measure that relies on the computationally intensive nonparametric bootstrap for confidence interval estimation. In the work described herein, we formulated a more efficient two-stage, cross-fitted estimation procedure for the R2 measure. To avoid potential bias, we performed iterative Sure Independence Screening (iSIS) in two subsamples to exclude the non-mediators, followed by ordinary least squares regressions for the variance estimation. We then constructed confidence intervals based on the newly derived closed-form asymptotic distribution of the R2 measure. Extensive simulation studies demonstrated that this proposed procedure is much more computationally efficient than the resampling-based method, with comparable coverage probability. Furthermore, when applied to the Framingham Heart Study, the proposed method replicated the established finding of gene expression mediating age-related variation in systolic blood pressure and identified the role of gene expression profiles in the relationship between sex and high-density lipoprotein cholesterol level. The proposed estimation procedure is implemented in R package CFR2M.
Stein, D. W.; Gaspar, F.; Piantadosi, S.; Amin, A.; Webb, B.; Lu, D.; D'Arinzo, L.; Oliver, M.; Fitzgerald, K.
Show abstract
Methods of causal inference are used to estimate treatment effectiveness for non-randomized study designs. The propensity score (i.e., the probability that a subject receives the study treatment conditioned on a set of variables related to treatment and/or outcome) is often used with matching or sample weighting techniques to, ideally, eliminate bias in the estimates of treatment effect due to treatment decisions. If multiple treatments are available, the propensity score is a function of the adjustment set and the set of possible treatments. This paper develops a compound model that separates the treatment decision into a binary decision: treat or dont treat; and a potential treatment decision: choose the treatment that would be given if the subject is treated. It is applicable if the treatment set is finite, treatments are given at one time point, and the outcome is observed at a fixed time point. This representation can reduce bias when not all treatments are available to all patients. Multiple treatment stabilized marginal structural weights were calculated with this approach, and the method was applied to an observational study to evaluate the effectiveness of different neutralizing monoclonal antibodies to treat infection with various severe acute respiratory syndrome coronavirus 2 variants.
Wu, D.; Goldfeld, K. S.; Petkova, E.; Park, H. G.
Show abstract
BackgroundPrecision medicine has led to the development of targeted treatment strategies tailored to individual patients based on their characteristics and disease manifestations. Although precision medicine often focuses on a single health outcome for individualized treatment decision rules (ITRs), relying only on a single outcome rather than all available outcomes information leads to suboptimal data usage when developing optimal ITRs. MethodsTo address this limitation, we propose a Bayesian multivariate hierarchical model that leverages the wealth of correlated health outcomes collected in clinical trials. The approach jointly models mixed types of correlated outcomes, facilitating the "borrowing of information" across the multivariate outcomes, and results in a more accurate estimation of heterogeneous treatment effects compared to using single regression models for each outcome. We develop a treatment benefit index, which quantifies the relative treatment benefit of the experimental treatment over the control treatment, based on the proposed multivariate outcome model. ResultsWe demonstrate the strengths of the proposed approach through extensive simulations and an application to an international Coronavirus Disease 2019 (COVID-19) treatment trial. Simulation results indicate that the proposed method reduces the occurrence of erroneous treatment decisions compared to a single regression model for a single health outcome. Additionally, the sensitivity analysis demonstrates the robustness of the model across various study scenarios. Application of the method to the COVID-19 trial exhibits improvements in estimating the individual-level treatment efficacy (indicated by narrower credible intervals for odds ratios) and optimal ITRs. ConclusionThe study jointly models mixed types of outcomes in the context of developing ITRs. By considering multiple health outcomes, the proposed approach can advance the development of more effective and reliable personalized treatment
Qian, Y.; Song, Y.
Show abstract
Instrumental variable (IV) methods are widely used in health and social sciences to estimate causal treatment effects among compliers. In certain research settings, the instrument-treatment association (first stage) and the instrument-outcome association (reduced form) are each estimated from a different dataset. Two-Sample Instrumental Variables (TSIV), proposed by Angrist and Krueger (1992), addresses this by combining first-stage and reduced-form estimates from separate data sources into a single causal effect estimate. However, TSIV identification requires that instrument compliance behavior be consistent across the two samples, a condition that is rarely verified in practice. We show mathematically and empirically that when compliance differs between samples, the raw TSIV estimator does not converge to the true Local Average Treatment Effect (LATE) and instead attenuates toward a predictably biased limit proportional to the ratio of first-stage compliance rates between the two samples. To address this, we formalize a framework for estimating LATE with TSIV under two key assumptions: (1) Covariate Overlap, requiring that the two samples share sufficient common support in their covariate distributions, and (2) Compliance Transportability, requiring that compliance behavior is identical across populations after conditioning on observed covariates. We consider a setting in which a health policy instrument and outcomes are recorded in administrative claims while treatment and covariates are collected in a survey. We use a C-statistic derived from pooled covariates to detect population mismatch and an Inverse Probability Weighting (IPW) correction that reweights the first-stage sample to approximate the administrative covariate distribution. In Monte Carlo simulations across eight scenarios calibrated to a survey-Medicaid setting, IPW-TSIV reduces bias in estimating the LATE, achieving 88% reduction in the primary scenario, 82% under severe selection, and 79% when state-level expansion policy drives compliance heterogeneity. We further validate this framework using the Oregon Health Insurance Experiment, where partitioning the public-use lottery data (N = 24,646) into two non-overlapping samples with substantively meaningful compliance heterogeneity yields a verifiable benchmark against the true causal effect. IPW-TSIV reduces mean absolute bias by 71.6% relative to the oracle S2-specific LATE across 10 independent replications (C-statistic = 0.78), outperforms naive TSIV in all 10 splits, and reduces mean bias relative to the full-data LATE from +0.016 to +0.008. This framework provides applied researchers with actionable diagnostic thresholds to detect sample mismatch, validate transportability assumptions, and determine when structural TSIV estimation is reliable.
Codi, A. M.; Kim, S.; Rogawski McQuade, E. T.; Benkeser, D.; AntiBiotics for Children with severe Diarrhea (ABCD) Study Group,
Show abstract
Acute diarrheal disease is one of the leading causes of death in children under age 5, disproportionately impacting children in low-resource settings. Many of these cases are caused by bacteria and therefore could respond to antibiotic treatment; however, the benefits of widely prescribing antibiotics must be weighed against the risks for the emergence of microbial resistance. These challenges present the opportunity for developing individualized treatment guidelines for diarrheal disease. In this study, we utilize a framework for the creation and evaluation of individualized treatment rules that leverage diagnostic and other clinical information to recommend antibiotic treatment to children with watery diarrhea. In contrast to many applications of pipelines for creating and evaluating treatment rules, we (i) explicitly consider creating rules that limit the proportion of children treated under a rule, to limit risks for overtreatment and the emergence of microbial resistance and (ii) propose methods to compare the performance of rules based on different sets of input covariates, which allows for quantification of the impact of measuring additional diagnostic biomarkers in clinical settings. We use a nested cross validation procedure that makes use of ensemble machine learning and doubly-robust estimation approach to derive, evaluate, and compare rules. We demonstrate that our proposed method yields appropriate inference in a realistic simulation study and apply our method to a real-data analysis of the AntiBiotics for Children with severe Diarrhea (ABCD) trial.
Oh, E. J.; Mikytuck, A.; Lancaster, V.; Goldstein, J.; Keller, S.
Show abstract
Understanding the prevalence of infections in the population of interest is critical for making data-driven public health responses to infectious disease outbreaks. Accurate prevalence estimates, however, can be difficult to calculate due to a combination of low population prevalence, imperfect diagnostic tests, and limited testing resources. In addition, strategies based on convenience samples that target only symptomatic or high-risk individuals will yield biased estimates of the population prevalence. We present Bayesian multilevel regression and poststratification models that incorporate probability sampling designs, the sensitivity and specificity of a diagnostic test, and specimen pooling to obtain unbiased prevalence estimates. These models easily incorporate all available prior information and can yield reasonable inferences even with very low base rates and limited testing resources. We examine the performance of these models with an extensive numerical study that varies the sampling design, sample size, true prevalence, and pool size. We also demonstrate the relative robustness of the models to key prior distribution assumptions via sensitivity analyses.
Radosavljevic, L.; Smith, S. M.; Nichols, T. E.
Show abstract
A particularly challenging form of missing data is structured missingness, where sets of subjects and variables consistently have missing data. For tabular data from sub-studies or modalities, structured missingness can come from non-participation in followup studies, which creates large blocks of missing data. Canonical Correlation Analysis (CCA) is a multivariate modelling tool commonly used to link two different set of variables, and in neuroimaging has typically been used to find associations between imaging and non-imaging variables. Motivated by CCA, we propose a new method for covariance estimation from incomplete data that handles data with a mix of structured and unstructured missingness, assuming Missing at Random (MAR). Our proposed method is compared to existing methodology by way of evaluation on simulated data and on real data from subjects in the UK Biobank brain imaging cohort.
Hazewinkel, A.-D.; Tilling, K.; Wade, K. H.; Palmer, T. M.
Show abstract
Randomized controlled trials (RCTs) are considered the gold standard for assessing the causal effect of an exposure on an outcome, but are vulnerable to bias from missing data. When outcomes are missing not at random (MNAR), estimates from complete case analysis (CCA) will be biased. There is no statistical test for distinguishing between outcomes missing at random (MAR) and MNAR, and current strategies rely on comparing dropout proportions and covariate distributions, and using auxiliary information to assess the likelihood of dropout being associated with the outcome. We propose using the observed variance difference across treatment groups as a tool for assessing the risk of dropout being MNAR. In an RCT, at randomization, the distributions of all covariates should be equal in the populations randomized to the intervention and control arms. Under the assumption of homogeneous treatment effects, the variance of the outcome will also be equal in the two populations over the course of followup. We show that under MAR dropout, the observed outcome variances, conditional on the variables included in the model, are equal across groups, while MNAR dropout may result in unequal variances. Consequently, unequal observed conditional group variances are an indicator of MNAR dropout and possible bias of the estimated treatment effect. Heterogeneity of treatment effect affects the intervention group variance, and is another potential cause of observing different outcome variances. We show that, for longitudinal data, we can isolate the effect of MNAR outcome-dependent dropout by considering the variance difference at baseline in the same set of patients that are observed at final follow-up. We illustrate our method in simulation and in applications using individual-level patient data and summary data.
Pan, W.; Lu, Z.; Jiang, W.; Lim, J.; Xu, L.; Wang, X.
Show abstract
In meta-analyses of continuous outcomes, the sample mean and standard deviation (SD) are essential for synthesizing effect sizes across studies. However, clinical studies frequently report alternative summary statistics, such as the median, quartiles, and range. To enable inclusion of such studies, various methods have been proposed to estimate the sample mean and SD from these reported summaries. We propose the Bayesian Order Statistics-based Estimator (BOSE), which leverages the joint likelihood of observed order statistics together with weakly informative priors to obtain the full posterior distribution for the mean and SD without relying on computationally intensive iterative procedures such as Markov chain Monte Carlo algorithms. Our numerical studies demonstrate that BOSE performs competitively with existing approaches in estimating the mean, while achieving superior performance for estimating the SD across all evaluated scenarios, particularly in small-sample settings. Under non-normal distributions including skewed, heavy-tailed, and bimodal settings with mild or moderate deviations from normality, BOSE remains robust and stable, whereas methods specifically designed for skewed distributions may become unstable or even inapplicable. Beyond point estimation, BOSE naturally provides empirically validated posterior credible intervals, enabling researchers to formally quantify uncertainty for study-level estimates and make reliable, evidence-based decisions in meta-analytic research synthesis. A publicly accessible web application implementing BOSE and competing methods is also provided to facilitate practical use in meta-analytic research.
Wang, H.; Zhang, B.; Lei, Y.; Lu, Y.; Zhang, D.; Jian, X.; Zhu, Y.; Hu, W.; Chu, H.; Chen, Y.; Suchard, M. A.; Ryan, P. B.; Hripcsak, G.; Asch, D. A.; Lu, Y.; Bin, Y.; Schuemie, M. J.; Qiu, Y.; Chen, Y.
Show abstract
Glucagon-like peptide-1 receptor agonists (GLP-1RAs) have been linked to heterogeneous, potentially pleiotropic effects across organ systems, motivating outcome-wide comparative risk profiling in real-world data. A central challenge in such analyses is \emph{residual bias} that remains after adjustment for observed confounders, which can distort effect estimates and mis-calibrate uncertainty. We present distributional diagnosis and calibration (DC), which uses panels of negative control outcomes (NCOs) to diagnose residual bias and calibrate uncertainty. DC evaluates null behavior via $p$-value uniformity and empirical coverage across NCOs, and uses the empirical distribution of NCO effect estimates to calibrate confidence intervals for prespecified primary outcomes. DC is modular: it can wrap around commonly used causal inference methods and operates directly on summary statistics, supporting collaborative research under data-sharing constraints. Using electronic health records from a large U.S. clinical research network (152.7 million patients), we compared GLP-1RAs with sodium--glucose cotransporter~2 inhibitors across 15 prespecified outcomes spanning cardiovascular, mental health, and genitourinary domains using four causal estimators. Across outcomes and methods, DC diagnostics revealed substantial and method-dependent residual systematic error. DC calibration attenuated systematic error signals observed in negative controls and yielded more stable, better-calibrated estimates for clinical outcomes, supporting DC as a practical strategy to strengthen the credibility of real-world comparative effectiveness research.
Schwenke, J. M.; Herkner, F.; Kayembe, M. T.; Olsen, I. C.; Briel, M.; König, F.
Show abstract
Acute viral respiratory infections (ARVIs) are a major cause of hospitalization and death worldwide, yet randomized clinical trials in this setting face substantial challenges in selecting efficient and clinically meaningful primary endpoints. Mortality is often too infrequent to serve as a feasible primary endpoint. Several alternative approaches have been proposed, including ordinal scales, time-to-event endpoints, recovery-based composite outcomes, and longitudinal ordinal models. However, their comparative operating characteristics under realistic ARVI disease courses remain insufficiently understood. We describe a simulation study to compare the type I error and power of commonly used and recently proposed endpoints and analysis strategies for two-arm randomized trials in hospitalized participants with ARVIs. Data will be generated under several mechanisms designed to mimic plausible participant trajectories, including a latent Brownian motion process, a first-order ordinal Markov process, a latent recurrent-event process with frailty, and resampling from individual participant data from the ACTT-2 trial. Simulated outcomes will use 4-, 6-, and 8-level ordinal severity scales and will reflect moderately and severely ill populations, follow-up horizons of 28 or 60 days, varying treatment effects, and sample sizes. Methods to be compared include Markov ordinal state transition models, proportional-odds models at a fixed time point, days-to-recovery scale analyses, Cox models for time-to-event endpoints, logistic regression for binary endpoints, generalized pairwise comparisons for hierarchical composites, and t-tests for days alive and out of hospital. This study will provide a systematic comparison of endpoint definitions and analysis methods for ARVI trials under clinically motivated data-generating mechanisms. The results are intended to inform the selection of feasible, interpretable, and statistically efficient primary analysis strategies for future trials in viral respiratory disease.
Irlmeier, R.; Jin, Z.; Ye, F.
Show abstract
Background Simon two-stage designs for binary endpoints and their time-to-event analogues, including the Kwak and Jung method, rely on a fixed null benchmark. Their Type I error control is valid only when that benchmark is correctly specified. In practice, historical benchmarks are often inconsistent due to small samples, population heterogeneity, changing eligibility criteria, and evolving standards of care. Even modest misspecifications can substantially inflate the Type I error rate, leading to costly advancement of ineffective treatments. Methods We propose the Interval-Null Robust (INR) two-stage design framework that accounts for uncertainty in the historical null benchmark. We define the null hypothesis as a plausible range of clinically uninteresting values: p[isin][p0L, p0U] for binary endpoints and {lambda}[isin][{lambda}0L, {lambda}0U] (or equivalent survival probabilities) for time-to-event endpoints. Type I error is controlled uniformly over the full null interval: sup{theta}[isin]{theta}0 Pr{theta}(Go) [≤] . Under the monotonicity of the Go probability, the supremum occurs at the least favorable null configuration - p0U and {lambda}0L - but the design is not reduced to a point-null formulation. The interval defines the uncertainty set for error control and is used in selecting among feasible designs through robust criteria such as worst-case regret or minimal average expected sample size. Results Across representative planning scenarios for both endpoint types, classic designs calibrated to a single benchmark exhibit substantial Type I error inflation when the true null parameter exceeds the assumed planning value. INR designs maintain the nominal Type I error rate across the full null interval, directly addressing this vulnerability to benchmark misspecification. The robustness-efficiency trade-off can be managed through design constraints and robust optimization criteria while preserving uniform Type I error control. Conclusions INR two-stage designs offer a transparent framework for addressing historical control uncertainty in single-arm Phase II trials. By replacing reliance on a fixed benchmark assumption with a more realistic interval of clinically plausible null values, INR design reduces the risk of false-positive Go-decisions caused by benchmark misspecification. INR applies to both binary and time-to-event endpoints and is implemented in the open-source INRDesign R package and accompanying interactive Shiny app.
Bollee, M.; Dutta Majumdar, A.
Show abstract
Discrete-time Markov cohort-state transition models are now well-established as the preferred choice of analysts across application areas including health technology assessment. This preference arises out of its relative intuition and its capability to strike a fine balance between complex disease pathways, statistical precision, and parsimony although being criticized by a wide variety of stakeholders. Transition probability matrices (TPMs) are the "heart and soul" of such models responsible for estimating patient dispositions. However, estimating such TPMs comes with its own set of challenges. In some situations, the transition data may be censored such that the health state of a patient is unknown for multiple time steps before the next observation or data immaturity especially in rare diseases. Craig and Sendi proposed the expectation-maximization (EM) algorithm using uniform weights as a solution for unequal estimation intervals for partially observed data. However, this typically comes at the cost of increased within-state output variations with no optimization technique available in the literature. The objective of this paper is to explore an optimized weighted version of the original EM algorithm, that aims to estimate the set of weights which minimizes the uncertainty of the estimated TPM against a target objective function. The weighting reduces the uncertainty of the estimate by considering the difference in temporal sparsity of the data when there are missing time steps. Further, we demonstrate the applicability of this weighting method using a fictitious cost-effectiveness model with our approach, showing a fine but definitive change over the original approach.
Hung, J.-Y.; Hsu, C.-Y.; Su, P.-F.; Shyr, Y.
Show abstract
A surrogate endpoint is a biomarker that is reasonably likely to predict clinical benefit and is used as a substitute for a direct measure of clinical benefit under the Food and Drug Administration (FDA) Accelerated Approval pathway. According to FDA guidelines, a valid surrogate endpoint must meet two associations: I-association (the association between the surrogate and true endpoints, such as disease response and overall survival) and T-association (the association between treatment effects on both endpoints, such as odds ratio and hazard ratio). I-association is commonly evaluated, but T-association is often overlooked due to the lack of appropriate statistical methods. Failure to satisfy T-association precludes a biomarker from supporting accelerated approval. To address this gap, we propose a new method to rigorously assess T-association in accordance with FDA guidelines. This method assumes that treatment effects on the surrogate and true endpoints follow a bivariate normal distribution, accounting for both within-study and between-study variances. The key evaluation metric is the correlation coefficient, which quantifies the relationship between treatment effects on both endpoints. Model parameters, including this correlation, are estimated using maximum likelihood, restricted maximum likelihood, and a Bayesian approach. We demonstrate the method using both simulated and real-world data. The method will serve as the statistical foundation that aligns with FDA guidelines and supports future accelerated approvals. The R package to implement the proposed method is available at https://github.com/jybelindahung/T-association.
Hegarty, S. E.; Linn, K. A.; Zhang, H.; Teeple, S.; Albert, P. S.; Parikh, R. B.; Courtright, K.; Kent, D. M.; Chen, J.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWThe proliferation of algorithm-assisted decision making has prompted calls for careful assessment of algorithm fairness. One popular fairness metric, equal opportunity, demands parity in true positive rates (TPRs) across different population subgroups. However, we highlight a critical but overlooked weakness in this measure: at a given decision threshold, TPRs vary when the underlying risk distribution varies across subgroups, even if the model equally captures the underlying risks. Failure to account for variations in risk distributions may lead to misleading conclusions on performance disparity. To address this issue, we introduce a novel metric called adjusted TPR (aTPR), which modifies subgroup-specific TPRs to reflect performance relative to the risk distribution in a common reference subgroup. Evaluating fairness using aTPRs promotes equal treatment for equal risk by reflecting whether individuals with similar underlying risks have similar opportunities of being identified as high risk by the model, regardless of subgroup membership. We demonstrate our method through numerical experiments that explore a range of differential calibration relationships and in a real-world data set that predicts 6-month mortality risk in an in-patient sample in order to increase timely referrals for palliative care consultations.
Shang, N.; Schrag, S.; Kahn, R.; Rhodes, J.
Show abstract
Establishing dose-response relationships from observational data is challenging due to confounding and sample selection bias. Standard causal methods adjust for confounding but typically require knowledge of covariate distributions in the target population--often via a well-defined probability sampling scheme. We propose the Covariate Adjusted Logit Model (CALM), which generalizes log-linear structural mean models for binary exposures to continuous exposures by modeling a relative dose-response curve anchored to a baseline level. By separating this curve from the null disease risk (NDR) at baseline, CALM enables valid inference under biased sampling while adjusting for confounding effects. A Gibbs sampler--the All-or-Nothing algorithm--is introduced to support Bayesian modeling, drawing on a vaccine-effect-inspired interpretation of the relative dose-response curve. Simulation studies demonstrate that CALM recovers dose-response relationships more accurately in the presence of bias and confounding. In vaccine trials, where confounding covariates affect immune responses differently across study arms, CALM provides a more accurate and robust antibody-disease curve to serve as a surrogate for evaluating vaccine effectiveness.
Joosten, R.; Abhishta, A.
Show abstract
We design a procedure (the complete Python code may be obtained at https://github.com/abhishta91/antibody_montecarlo) using Monte Carlo (MC) simulation to establish the point estimators described below and confidence intervals for the base rate of occurence of an attribute (e.g., antibodies against Covid-19) in an aggregate population (e.g., medical care workers) based on a test. The requirements for the procedure are the tests sample size (N) and total number of positives (X), and the data on tests reliability. The modus is the prior which generates the largest frequency of observations in the MC simulation with precisely the number of test positives (maximum-likelihood estimator). The median is the upper bound of the set of priors accounting for half of the total relevant observations in the MC simulation with numbers of positives identical to the tests number of positives. O_LSTOur rather preliminary findings areC_LSTO_LIThe median and the confidence intervals suffice universally. C_LIO_LIThe estimator [Formula] may be outside of the two-sided 95% confidence interval. C_LIO_LIConditions such that the modus, the median and another promising estimator which takes the reliability of the test into account, are quite close. C_LIO_LIConditions such that the modus and the latter estimator must be regarded as logically inconsistent. C_LIO_LIConditions inducing rankings among various estimators relevant for issues concerning over-or underestimation. C_LI JEL-codes: C11, C13, C63