Back

Standardizing COVID-19 surveillance data into the OMOP common data model: a first implementation case study from Senegal

Diop, O.; Odhiambo, R.; Diouf, O.; Momanyi, R.; Ochola, M.; Diallo, A. S.; Padane, A.; Cygu, S. B.; Barasa, M.; Iddi, S.; Kiragga, A.; Sarr, M.; Mboup, S.; Mboup, A.

2026-07-06 health informatics
10.64898/2026.07.01.26357078 medRxiv
Show abstract

The COVID-19 pandemic highlighted the need for interoperable health data infrastructures supporting reproducible observational research. The Observational Medical Outcomes Partnership Common Data Model (OMOP CDM) provides a widely adopted standard for harmonizing heterogeneous health data, but adoption remains limited in francophone Africa where language barriers and non-standardized surveillance systems pose additional challenges. We developed a complete Extract-Transform-Load (ETL) pipeline to convert a heterogeneous Senegalese COVID-19 surveillance dataset into OMOP CDM version 5.4. Source data recorded in French were translated into English through an iterative process interleaved with vocabulary mapping using ATHENA and Usagi. Semantic standardization used SNOMED CT for conditions, LOINC for measurements, and RxNorm for drugs. All 214 mappings underwent expert review by clinical and data science specialists. Data quality was assessed using the OHDSI Data Quality Dashboard (DQD) and Achilles. The standardized database achieved complete transformation (100%) for eight of the eleven source-populated domain tables, including person, visit_occurrence, measurement, and death. Partial transformation was observed for condition_occurrence (95.3%) and observation (68.1%), primarily due to incomplete vocabulary coverage for occupation categories and context-specific variables. The DQD produced an overall pass rate of 97% and a corrected pass rate of 98%, comparable to other published African OMOP implementations. Among the 19 data-quality failures, conformance and completeness issues predominated; the conformance failures were largely foreign-key checks, reflecting placeholder concept values (concept_id = 0) for metadata fields without meaningful equivalents in surveillance data. Iterative translation refinement was required when French-to-English translations did not align with OHDSI vocabulary terminology. This work documents, to our knowledge, the first OMOP CDM implementation on COVID-19 surveillance data in Senegal and francophone West Africa and provides a reusable methodological blueprint for future OMOP deployments in the region.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Scientific Data
209 papers in training set
Top 0.2%
10.8%
2
PLOS ONE
5266 papers in training set
Top 20%
8.7%
3
GigaScience
212 papers in training set
Top 0.2%
8.7%
4
PLOS Digital Health
106 papers in training set
Top 0.7%
7.7%
5
JMIR Medical Informatics
18 papers in training set
Top 0.1%
7.7%
6
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.3%
6.6%
50% of probability mass above
7
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.6%
5.4%
8
BMC Medical Research Methodology
47 papers in training set
Top 0.2%
5.4%
9
JAMIA Open
42 papers in training set
Top 0.5%
4.0%
10
Scientific Reports
3612 papers in training set
Top 35%
3.2%
11
DIGITAL HEALTH
17 papers in training set
Top 0.3%
2.3%
12
BMJ Open
601 papers in training set
Top 8%
2.3%
13
Nature Communications
5641 papers in training set
Top 42%
2.1%
14
BMJ Health & Care Informatics
15 papers in training set
Top 0.5%
1.9%
15
Journal of Medical Internet Research
87 papers in training set
Top 2%
1.7%
16
International Journal of Medical Informatics
26 papers in training set
Top 0.7%
1.7%
17
Frontiers in Digital Health
24 papers in training set
Top 0.8%
1.7%
18
Wellcome Open Research
67 papers in training set
Top 0.9%
1.3%
19
Communications Medicine
113 papers in training set
Top 3%
1.3%
20
Journal of Biomedical Informatics
47 papers in training set
Top 1%
1.0%
21
The Lancet Digital Health
25 papers in training set
Top 0.7%
0.8%
22
BMJ
51 papers in training set
Top 1%
0.8%
23
Database
61 papers in training set
Top 1%
0.8%
24
Orphanet Journal of Rare Diseases
21 papers in training set
Top 0.6%
0.8%
25
Frontiers in Public Health
148 papers in training set
Top 6%
0.8%