Data processing pipelines and tools for routine health facility malaria surveillance in Uganda
Carter, A. R.; Smith, D. L.; Enganyu, T.; Hergott, D. E. B.; Rek, J.; Okiring, J.; Kyabayinze, D.; Maiteki, C.; Mbaka, P.
Show abstract
Timely processing of routine health facility surveillance data is essential for responsive malaria control, yet the path from raw electronic reports to actionable intelligence remains a largely undocumented challenge in endemic countries. We present an open-source Extract-Transform-Load (ETL) software system and accompanying metadata R package (ramptools) that together convert raw DHIS2 health facility data into cleaned, version-controlled, analysis-ready datasets for Uganda's National Malaria Elimination Division (NMED). The ETL pipeline is implemented in R and deployed on a cloud platform with scripts running on automated schedules to keep the database current. The system extracts malaria indicators from two successive DHIS2 instances via the DHIS2 Web API, applies a two-stage outlier detection algorithm combining variance-based screening with STL decomposition, imputes missing values through seasonal interpolation, enforces logical consistency constraints across indicator cascades, and aggregates facility-level data through a six-level administrative hierarchy. All raw data are stored in an append-only versioned schema, enabling reconstruction of the database state at any historical point. The ramptools package provides standardized metadata - including an indicator crosswalk mapping 118 indicators across DHIS2 instances, a location hierarchy of 11,229 organizational units, geolocated health facility attributes, and administrative boundary shapefiles - as lazy-loaded R data objects. We validate the system through a proof-of-principle outbreak detection application that consumes ETL outputs to compute district-level outbreak indices via kernel-smoothed time series analysis, deployed as an interactive Shiny dashboard. The software has been in development since 2020, processing weekly and monthly data for approximately 8,700 health facilities. All code is open source (MIT license) and hosted on GitHub.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Geographic Name Resolution Service: A tool for the standardization and indexing of world political division names, with applications to species distribution modeling 94%
- easyFulcrum: An R package to process and analyze ecological sampling data generated using the Fulcrum mobile application 94%
- Extract, Transform, Load Framework for the Conversion of Health Databases to OMOP 93%
Similar papers in this journal
- The Role of Modelling and Analytics in South African COVID-19 Planning and Budgeting 92%
- Coaching visits and supportive supervision for primary care facilities to improve malaria service data quality in Ghana: an intervention case study 91%
- Estimating the relationship between mobility, non-pharmaceutical interventions, and COVID-19 transmission in Ghana 90%
Similar papers in this journal
- Electronic data capture for large scale typhoid surveillance, household contact tracing, and health utilisation survey: Strategic Typhoid Alliance across Africa and Asia. 91%
- Baseline nowcasting methods for handling delays in epidemiological data 90%
- Re-examining the meningitis belt: associations between environmental factors and epidemic meningitis risk across Africa 89%
Similar papers in this journal
- The MARC SE-Africa Dashboard: Joining Forces to Counteract Emerging Antimalarial Resistance in South and East Africa 94%
- COVID-19 Vaccination Data Management and Visualization Systems for Improved Decision-Making: Lessons Learnt from Africa CDC Saving Lives and Livelihoods Program 93%
- From months to minutes: creating Hyperion, a novel data management system expediting data insights for oncology research and patient care 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.