Back

MedZone Embedder: a framework for representation learning of Japanese secondary medical care areas from a national ICU registry, characterizing intensive care provision structure and regional vulnerability

Ohno, K.; Hashimoto, S.

2026-07-20 health informatics
10.64898/2026.07.17.26358373 medRxiv
Show abstract

Background: In Japan, acute inpatient care is divided into approximately 335 secondary medical care areas, which serve as the basic units for planning healthcare delivery systems under the 8th National Health Care Plan. While comparisons between regions and facilities typically rely on a single risk-adjusted metric, this approach confuses differences in patient demographics with differences in the actual infrastructure of intensive care units (ICUs). This paper presents a framework - MedZone Embedder - for deriving data-driven indicators of regional structural vulnerability by mapping secondary medical care areas onto a learned similarity space, together with its working implementation. The paper sets out the concept, the method, a proof of concept, and an explicit staged validation program, rather than national empirical results. Methods: Each area is represented by a feature vector consisting of aggregated values of intensive care provision indicators derived directly from the Japan Intensive Care Patient Database (JIPAD) - specifically, risk-adjusted mortality rates (standardized mortality ratios and an in-hospital composite indicator), technical efficiency, length of stay, readmission rates, case severity, and case composition - with the within-area variance of these indicators also taken into account. No hierarchical processing by facility type is performed. A contrastive autoencoder (multilayer perceptron encoder 32 -> 16 -> 8, symmetric decoder) is trained by self-supervised learning, using an objective function that combines reconstruction and normalized temperature cross-entropy (NT-Xent) on noise-augmented views. The resulting 8-dimensional embedding supports area searches based on cosine similarity and anomaly scoring in the embedding space (using isolation forest, Mahalanobis distance, or k-nearest-neighbor density), which is normalized to a vulnerability score ranging from 0 to 1. If deep learning libraries are unavailable, or if the number of areas is small, an alternative method using deterministic principal component analysis is employed. Results: This method was implemented and deployed within an operational ICU decision support system on a managed cloud platform. The proof of concept (PoC) is structured around five secondary medical care areas within Kyoto Prefecture and runs entirely on synthetic facility-level aggregate data constructed to follow the JIPAD indicator schema; no registry data were accessed. It generated: an aggregate provision profile for each area; an area embedding space equipped with a similar-area search function; and a vulnerability ranking that identifies areas with low patient numbers and low diversity that exhibit overall poor outcomes. At this scale, the contrastive autoencoder falls back to principal component projection. The deep learning pathway has been implemented and unit testing has been completed; training and evaluation on actual registry data are pending data-use approval and the expansion of data integration. Validation is staged: Stage 2 will train the contrastive pathway over JIPAD-covered areas to assess construct validity against public structural indicators (ICU/HCU beds, population, accessibility), and Stage 3 will extend coverage to all areas via National Database (NDB) linkage. Conclusion: MedZone Embedder reframes regional comparison from single-indicator ranking to structural representation: which areas are alike, and which are structural outliers. The contribution of this paper is the framework - the proposal that the intensive care provision structure of Japanese secondary medical care areas can be learned from a national outcomes registry and read through the lens of what we call institutional debt - together with a deployed implementation and a pre-specified validation program. To our knowledge, this is a candidate first application of contrastive representation learning to Japanese secondary medical care areas.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.1%
14.7%
2
Scientific Reports
3612 papers in training set
Top 9%
7.1%
3
PLOS Digital Health
106 papers in training set
Top 0.9%
6.1%
4
International Journal of Medical Informatics
26 papers in training set
Top 0.2%
4.7%
5
JMIR Medical Informatics
18 papers in training set
Top 0.1%
4.7%
6
Frontiers in Digital Health
24 papers in training set
Top 0.2%
4.7%
7
Journal of Medical Internet Research
87 papers in training set
Top 0.5%
4.7%
8
PLOS ONE
5266 papers in training set
Top 32%
4.7%
50% of probability mass above
9
Patterns
78 papers in training set
Top 0.4%
3.9%
10
IEEE Journal of Biomedical and Health Informatics
37 papers in training set
Top 0.3%
3.9%
11
Computers in Biology and Medicine
128 papers in training set
Top 1%
3.1%
12
Frontiers in Public Health
148 papers in training set
Top 2%
2.6%
13
JMIR Public Health and Surveillance
45 papers in training set
Top 0.3%
2.6%
14
GigaScience
212 papers in training set
Top 2%
2.3%
15
Artificial Intelligence in Medicine
17 papers in training set
Top 0.3%
2.1%
16
npj Digital Medicine
118 papers in training set
Top 2%
1.9%
17
Nature Communications
5641 papers in training set
Top 46%
1.7%
18
Journal of Biomedical Informatics
47 papers in training set
Top 0.8%
1.6%
19
Expert Systems with Applications
11 papers in training set
Top 0.2%
1.5%
20
IEEE Access
35 papers in training set
Top 0.9%
1.3%
21
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
1.3%
22
DIGITAL HEALTH
17 papers in training set
Top 0.7%
1.1%
23
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 0.9%
1.0%
24
Computational and Structural Biotechnology Journal
242 papers in training set
Top 6%
1.0%
25
The Lancet Digital Health
25 papers in training set
Top 0.8%
0.8%
26
iScience
1154 papers in training set
Top 37%
0.8%
27
Viruses
332 papers in training set
Top 5%
0.8%
28
JAMIA Open
42 papers in training set
Top 2%
0.8%
29
Wellcome Open Research
67 papers in training set
Top 2%
0.8%
30
Life
29 papers in training set
Top 1%
0.6%