Back

The Synthetic Fidelity-Stability Framework (SFSF): A Systematic Multi-Dimensional Benchmark of Synthetic Clinical Laboratory Data Generators

Desh, S. S.; Achary, P. M.; Nayak, S.

2026-08-12 biochemistry
10.64898/2026.08.11.741471 bioRxiv
Show abstract

BackgroundSynthetic data generation is increasingly proposed as a strategy to support privacy-preserving data sharing, augmentation of small or restricted biomedical datasets, and benchmarking of artificial intelligence tools in laboratory medicine. However, model selection remains difficult because synthetic data generators differ in fidelity, privacy risk, stability, and generalisability. Existing evaluations have rarely examined performance jointly across conditioning signal strength, synthetic output scale, and train-test generalisation. MethodsWe developed the Synthetic Fidelity-Stability Framework (SFSF), a systematic benchmark of 17 synthetic tabular data generation models using NHANES as a complex biomedical reference dataset. Models included statistical, copula-based, resampling, variational autoencoder, generative adversarial network, and diffusion-based approaches. Synthetic datasets were generated across 11 seed sizes, from 0 to 500 real conditioning observations, and six output scales, from 50 to 5,000 rows, yielding 1,122 synthetic datasets per run. Each dataset was evaluated against the full original dataset, the training subset, and a held-out test subset across five tiers: univariate distributional fidelity, moment agreement, tail behaviour, multivariate dependency structure, and privacy/memorisation risk. Composite rankings and seed-versus-output stability profiles were derived. ResultsUnivariate fidelity was broadly recovered across model classes and was the least discriminating tier. Resampling-based methods ranked highest overall but showed the greatest privacy risk, reflecting proximity to real observations rather than true generative novelty. VAE-family models reproduced moment statistics relatively well but consistently failed on tail and shape fidelity. GAN-family models showed substantial moment-level instability, while VineCopula demonstrated severe multivariate dependency failure. Diffusion-based models, particularly ForestDiffusion, provided the most favourable privacy-utility balance, combining competitive fidelity with the lowest privacy risk and the smallest train-test gap. ConclusionsNo single synthetic data generator dominated across fidelity, stability, and privacy dimensions. The SFSF framework provides a practical, multi-criterion approach for selecting synthetic tabular data generators according to intended clinical laboratory use, balancing statistical realism, dependency preservation, privacy risk, and robustness to seed and output scale.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.