TDA Engine v2.1: A Computational Framework for Detecting Structural Voids in Spatially Censored Epidemiological Data with Temporal Classification and Causal Inference
Mboya, G. O.
Show abstract
BackgroundIn public health surveillance, silence--the absence of data--is often more significant than the signal. Traditional epidemiological mapping tools efficiently visualize data density but struggle to mathematically define data absence. Standard approaches conflate stochastic sparsity with systemic suppression and remain vulnerable to edge effects. MethodsWe introduce a topological framework that detects structural voids--regions of unexpected data absence within clusters. Using Distance-to-Measure (DTM) filtration with adaptive thresholding via the Kneedle algorithm [11], we eliminate arbitrary parameter choices. Version 2.1 extends the original framework with three methodological additions: (1) a temporal void classifier combining the Fano factor and a two-state Hidden Markov Model (HMM) to distinguish persistent structural silence from stochastic fluctuation across reporting periods; (2) a causal taxonomy (BORDER, ACCESS, INFRASTRUCTURE, SYSTEM, UNKNOWN) that maps detected voids to probable reporting failure mechanisms via covariate decision trees; and (3) an Observed-to-Expected (O/E) completeness engine calibrated against WHO-standard disease incidence rates across seven conditions. Parameters are derived geometrically from the DTM distribution itself. We validate against known ground truth through a censoring simulation framework using public Kenyan health facility data. Detection accuracy is quantified using the Jaccard index [12], centroid error, and recovery rate. ResultsTDA Engine achieves Jaccard = 0.82 (95% CI: 0.74-0.89) on simulated suppression events, significantly outperforming KDE (0.45) and relative risk surfaces (0.38). Centroid error is 342 m (IQR: 187-512 m). The temporal classifier correctly labels 91% of structurally silent units across six-period validation datasets (HMM posterior P (structural) [≥]0.60). Permutation tests yield p = 0.003 (95% CI: 0.001-0.008) [13], confirming statistical significance beyond complete spatial randomness. ConclusionTDA Engine v2.1 provides a mathematically rigorous, topology-based framework for detecting structural voids in censored epidemiological data and classifying them by temporal persistence and probable causal mechanism. By shifting from density-based to geometry-based inference with quantitative validation metrics and causal labelling, we enable public health officials to distinguish between natural gaps and potential suppression, and to direct field investigation resources accordingly. We emphasize that structural voids are geometric anomalies consistent with suppression, not proof thereof--requiring contextual validation.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Flexible Framework for Local-Level Estimation of the Effective Reproductive Number in Geographic Regions with Sparse Data 94%
- Data-Driven Prediction of COVID-19 Cases in Germany for Decision Making 92%
- COVID19-Global: A shiny application to perform a global comparative data visualization for the SARS-CoV-2 epidemic 91%
Similar papers in this journal
- Accessibility of covariance information creates vulnerability in Federated Learning frameworks 91%
- spread.gl: visualising pathogen dispersal in a high-performance browser application 91%
- Interactive network-based clustering and investigation of multimorbidity association matrices with associationSubgraphs 91%
Similar papers in this journal
- A Bayesian Monte Carlo approach for predicting the spread of infectious diseases 94%
- Measuring the accuracy of gridded human population density surfaces: a case study in Bioko Island, Equatorial Guinea 93%
- Using mobile phone data to estimate dynamic population changes and improve the understanding of a pandemic: A case study in Andorra 93%
Similar papers in this journal
- Combining participatory mapping and route optimization algorithms to inform the delivery of community health interventions at the last mile 93%
- Emulation of epidemics via Bluetooth-based virtual safe virus spread: experimental setup, software, and data 91%
- COVID-19 Vaccination Data Management and Visualization Systems for Improved Decision-Making: Lessons Learnt from Africa CDC Saving Lives and Livelihoods Program 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.