CohortSymmetry: An R package to perform sequence symmetry analysis using the OMOP common data model
Chen, X.; Stanford, T.; Guo, Y.; Raventos, B.; Du, M.; Li, X.; Lam, A.; Corby, G.; Mercade-Besora, N.; Alcalde Herraiz, M.; Lopez-Guell, K.; Delmestri, A.; Man, W. Y.; PRIETO-ALHAMBRA, D.; Burn, E.; Catala, M.; Pratt, N.; Jodicke, A.; Newby, D.
Show abstract
BackgroundReal-world data are valuable for detecting adverse drug events, and Sequence Symmetry Analysis (SSA) is a simple yet effective method frequently used for this purpose. However, heterogeneous implementations across studies limit reproducibility and scalability. To address this, we developed an open-source R package that standardises SSA analytics using data mapped to the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM). MethodsWe developed CohortSymmetry, an R package that implements SSA for OMOP CDM data. The package was validated through unit testing and evaluated empirically by estimating adjusted sequence ratios (ASRs) with 95% confidence intervals (CIs) for 23 positive and 10 negative controls across six European databases, including CPRD GOLD (UK) and THIN(R) (Belgium, Italy, Romania, Spain, UK). Sensitivity and specificity were defined as the proportions of positive and negative controls correctly identified by SSA. Sensitivity analyses varied key parameters, including the washout period. ResultsCohortSymmetry passed high-coverage unit tests. Of 33 eligible controls, four showed results consistent with expectations across all databases; for example, the amiodarone-levothyroxine pair had a lower 95% CI bound >1 in each. Sensitivity was moderate, whereas specificity was high in the primary analyses. Parameter variation influenced outcomes; a 365-day prior observation requirement reduced specificity in CPRD GOLD from 75% to 38%. ConclusionsCohortSymmetry enables reproducible SSA using OMOP CDM data. Differences across databases likely reflect heterogeneity in data capture and prescribing patterns. Limitations include residual data variability and SSAs susceptibility to time-varying confounding, underscoring the need for tailored analytic design in pharmacovigilance studies. Key MessagesO_LIWe developed CohortSymmetry, an open-source R package that standardises SSA analytics using OMOP CDM-mapped data and verified the correctness of functions via unit testing and application to real-world datasets. C_LIO_LICohortSymmetry passed high-coverage tests, and among 33 selected controls, four showed results consistent with expectations across all databases; varying analytical parameters affected results. C_LIO_LIThe package provides a reproducible and scalable framework for multi-database SSA studies, supporting robust pharmacovigilance, but careful specification of parameters is required to account for the characteristics of the medical domain under investigation. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- People exposed to proton pump inhibitors shortly preceding COVID-19 diagnosis are not at an increased risk of subsequent hospitalizations and mortality: a nation-wide matched cohort study 94%
- Calcium Channel Blockers: clinical outcome associations with reported pharmacogenetics variants in 32,000 patients 92%
- Statin treatment effectiveness and the SLCO1B1 *5 reduced function genotype: long-term outcomes in women and men 90%
Similar papers in this journal
- Using quantitative bias analysis to adjust for misclassification of COVID-19 outcomes: An applied example of inhaled corticosteroids and COVID-19 outcomes 92%
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 92%
- Bias amplification of unobserved confounding in pharmacoepidemiological studies using indication-based sampling: there is no free lunch in restricting the sample to those with a particular drug-indication 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.