A Comprehensive Approach to Days' Supply Estimation in a Real-World Prescription Database: Data Cleaning, Imputation, and Adherence Analysis
Malk, M.; Mooses, K.; Oja, M.; Holm, J.; Keidong, H.; Umov, N.; Tamm, S.; Reisberg, S.; Vilo, J.; Kolde, R.
Show abstract
BackgroundFor accurate medication usage statistics and medication adherence calculations, we need to have an accurate days supply (DS) for each prescription. Unfortunately, often the DS or information needed for calculating the DS is not provided. Therefore, other methods need to be applied to acquire missing values or substituting incorrect values. ObjectiveThe aim of this study is to apply a variety of methods for managing incomplete and missing data to enhance the accuracy of calculating DS for all medications and drug forms alike. Furthermore, to describe the effect of applied methods on the medication adherence calculated on real-world data. MethodsA dataset comprising prescription records from a 10% random sample of the Estonian population between 2012 and 2019 was used. The workflow consisted of three steps - data cleaning, imputation and calculation of DS. For imputation, different methods were combined, such as calculating mode-based daily dose, or using usage guidelines from Summary of Product Characteristics (SPCs) or legislation. DS was calculated based on provided daily dose or imputed value. To evaluate the impact of data cleaning, medication adherence for baseline dataset and corrected dataset for two time periods 2012-2015 and 2017-2019 was calculated and compared. ResultsThe drug forms with the lowest proportion of correct DS provided were insulin injections (3.1%) and intravaginal contraceptives (8.0%) while the highest proportion of DS was provided for inhalation medication (57.5%), oral drops (53.0%) and tablets, capsules, suppositories (45.8%). As a result of applying different imputation approaches, we successfully found the DS for 98.3% (N=7,415,347) of dispensed prescriptions. For the remaining 1.7% (N=129,545) of prescriptions DS could not be imputed nor calculated with these methods. As for the medication adherence, the distinction between two observed time periods was more distinct in the baseline dataset compared with the corrected dataset for most of the drug groups, indicating that the applied correction methods had lessened the stark contrast. ConclusionsIn summary, our study demonstrated that with a carefully designed imputation pipeline where data-driven imputation is combined with domain knowledge and literature information, it is possible to meaningfully improve the quality of prescription datasets and generate more accurate and consistent adherence metrics across various drug form. Nonetheless, future efforts should continue to refine imputation techniques, incorporate machine learning approaches where appropriate, and expand validation efforts using external benchmarks or clinical outcomes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 95%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 93%
- Trajectories: a framework for detecting temporal clinical event sequences from health data standardized to the OMOP Common Data Model 92%
Similar papers in this journal
- Impact of CancelRx on Discontinuation of Controlled Substance Prescriptions 93%
- Development of a Mobile Application to Represent Food Intake in Inpatients: Dietary Data Systematization 90%
- Explainable AI enables clinical trial patient selection to retrospectively improve treatment effects in schizophrenia 90%
Similar papers in this journal
- Extent and causes of the collapse in the registration of innovative medications in Lebanon: A mixed-methods analysis 94%
- Assessment of knowledge and perception of prescribers towards rational medicine use in the Ashanti Region of Ghana 94%
- Adverse drug reactions associated with the use of biological agents 93%
Similar papers in this journal
Similar papers in this journal
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 92%
- Model-based reasoning methods for diagnosis in integrative medicine based on electronic medical records and natural language processing 92%
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.