Towards a Physiological Scaling Law: Model Quality vs. Cohort Size for Stochastic Sequence Data
Sunil, G.; Kumar, B. R.; Ramsundar, B.; Subramanian, S.
Show abstract
Scaling laws help determine the optimal data size for training large models but are established in domains where the target is deterministic. Physiological signals are different: heartbeat sequences are stochastic, so part of the error is irreducible even with large amounts of data. Metrics such as MAE do not account for non-deterministic behavior, and therefore assessing scaling requires evaluating distributional calibration (measuring how well predicted probability densities capture true conditional characteristics). We formulate a scaling law metric(n) = E + A n- and evaluate it with five metrics: accuracy (MAE, RMSE), distributional calibration (KS distance, goodness-of-fit), and training objective (negative log loss) using a neural temporal point process trained on a cohort of four-ECG datasets. The law fits all five metrics. While point accuracy is near saturation at n = 183, KS distance and goodness-of-fit improve by 6% and 12% respectively when extrapolated to 10,000 subjects, showing that scaling decisions in stochastic domains must be guided by distributional calibration rather than point accuracy.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A scaling approach to estimate the COVID-19 infection fatality ratio from incomplete data 92%
- Rett syndrome severity estimation with the BioStamp nPoint using interactions between heart rate variability and body movement 91%
- DeepGANnel: Synthesis of fully annotated single molecule patch-clamp data using generative adversarial networks 91%
Similar papers in this journal
- Average beta burst duration profiles provide a signature of dynamical changes between the ON and OFF medication states in Parkinson's disease 93%
- Inferring a simple mechanism for alpha-blocking by fitting a neural population model to EEG spectra 92%
- Exploring neural manifolds across a wide range of intrinsic dimensions 92%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.