DBToken: A Database Tokenizer for Medical Event Foundation Models
Shin, I.; McCann, K.; Marino, G.; Siam, U. T.; Li, H.; Stutz, E.; Edara, R.; Loza, A. J.
Show abstract
Objectives Transformer models for electronic health records require converting clinical data into token sequences, however standardized tokenization and evaluation frameworks are lacking. We introduce DBToken, an open-source library, and bits-per-row (BPR), a metric for comparing tokenization strategies. Materials and Methods DBToken accepts Medical Event Data Standard (MEDS)-compatible input and supports multiple text, numeric, and temporal tokenization strategies. BPR extends the bits-per-byte metric used in language models to enable comparison across tokenization strategies. Results DBToken efficiently tokenized data across configurations. BPR identified the vocabulary size associated with the best clinical outcome performance and localized differences in numeric tokenization performance by token class. Discussion Optimal tokenization strategies for medical foundation models are a subject of active research. DBToken enables reproducible tokenization experiments, while BPR efficiently screens vocabulary sizes and numeric representations before downstream evaluation. Conclusion DBToken and the BPR metric provide open-source infrastructure for reproducible EHR tokenization and cross-strategy evaluation.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review 95%
- Real World Performance of the 21st Century Cures Act Population Level Application Programming Interface 95%
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 94%
Similar papers in this journal
Similar papers in this journal
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 97%
- ConceptWAS: a high-throughput method for early identification of COVID-19 presenting symptoms 93%
- Automatic phenotyping of electronical health record: PheVis algorithm 93%
Similar papers in this journal
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 95%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 93%
- Extracting social determinants of health from electronic health records: development and comparison of rule-based and large language models-based methods 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.