IEEE Journal of Biomedical and Health Informatics
● Institute of Electrical and Electronics Engineers (IEEE)
All preprints, ranked by how well they match IEEE Journal of Biomedical and Health Informatics's content profile, based on 37 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Liu, M.; Dong, Q.; Wang, C.; Cheng, X.; Febrinanto, F. G.; Hoshyar, A. N.; Xia, F.
Show abstract
The wide variation in symptoms of neurological disorders among patients necessitates uncovering individual pathologies for accurate clinical diagnosis and treatment. Current methods attempt to generalize specific biomarkers to explain individual pathology, but they often lack analysis of the underlying pathogenic mechanisms, leading to biased biomarkers and unreliable diagnoses. To address this issue, we propose a motif-induced subgraph generative learning model (MSGL), which provides multi-tiered biomarkers and facilitates explainable diagnoses of neurological disorders. MSGL uncovers underlying pathogenic mechanisms by exploring representative connectivity patterns within brain net-works, offering motif-level biomarkers to tackle the challenge of clinical heterogeneity. Furthermore, it utilizes motif-induced information to generate enhanced brain network subgraphs as personalized biomarkers for identifying individual pathology. Experimental results demonstrate that MSGL outperforms baseline models. The identified biomarkers align with recent neuroscientific findings, enhancing their clinical applicability.
Deng, G.; Niu, M.; Luo, Y.; Rao, S.; Sun, J.; Xie, J.; Yu, Z.; Liu, W.; Zhao, S.; Pan, G.; Li, X.; Deng, W.; Guo, W.; Li, T.; Jiang, H.
Show abstract
Sleep disorders affect billions worldwide, yet clinical polysomnography (PSG) analysis remains hindered by labor-intensive manual scoring and limited generalizability of automated sleep staging tools across heterogeneous protocols. We present LPSGM, a large-scale PSG model designed to address two critical challenges in sleep medicine: cross-center generalization and adaptable diagnosis of neuropsychiatric disorders. Trained on 220,500 hours of multi-center PSG data (24,000 full-night recordings from 16 public datasets), LPSGM integrates domain-adaptive pre-training, flexible channel configurations, and a unified architecture to mitigate variability in equipment, montages, and populations during sleep staging while enabling downstream fine-tuning for brain disorder detection. In prospective validation, LPSGM achieves expert-level consensus in sleep staging ({kappa} = 0.845 {+/-} 0.066 vs. inter-expert {kappa} = 0.850 {+/-} 0.102) and matches the performance of fully supervised models on two independent private cohorts. When fine-tuned for sleep disorder diagnosis, LPSGM achieved 80.47% accuracy on the large-scale MNC dataset (773 subjects) for a three-class classification (Healthy Control vs. T1 Narcolepsy vs. Other Hypersomnia). The model also demonstrated strong cross-institutional generalizability, with an AUC of 0.8791 on independent cohorts for a binary (Normal vs. Abnormal) classification. While depression screening on smaller datasets showed perfect accuracy in controlled settings, larger-scale validation is necessary. By bridging automated sleep staging with real-world clinical deployment, LPSGM establishes a scalable framework for integrated sleep and brain disorder diagnostics. The code and pre-trained model are publicly available at https://github.com/Deng-GuiFeng/LPSGM to advance reproducibility and translational research in sleep medicine.
Kurt, F.; Subasi, S. N.; Yakisan, E. S.; Subasi, A.
Show abstract
Background: Wearable technologies enable scalable and continuous monitoring of emotional states through passive sensing of physiological and behavioral signals. However, conventional learning approaches often struggle to model the complex temporal, contextual, and relational dependencies underlying human emotions. To address these limitations, we propose a graph-based framework that represents multimodal wearable observations as heterogeneous knowledge graphs enriched with semantic information derived from Large Language Models (LLMs), enabling richer contextual understanding beyond raw sensor measurements. Methods: We constructed a heterogeneous knowledge graph using multimodal Fitbit physiological signals and affective self-report data collected from 45 users. Framing mood prediction and emotion detection was formulated as both binary and ternary node classification tasks. We evaluated five baseline heterogeneous Graph Neural Network (GNN) architectures and compared them with the proposed Semantically Gated Augmented Graph Neural Network (SeGA-GNN) framework, which dynamically integrates LLM-generated semantic embeddings into graph representations through a gated cross-modal fusion mechanism. Results: The baseline GNN models achieved strong performance, with classification accuracies ranging from 0.7525 to 0.9739 for binary classification and 0.6249 to 0.9699 for ternary classification. The proposed SeGA framework consistently improved predictive performance across most architectures. In particular, semantic augmentation transformed the HAN model from moderate baseline performance into near-perfect emotion recognition capability, achieving SeGA-HAN Accuracy = 0.9988 and AUC = 1.0000 for binary classification and Accuracy = 0.9979 and AUC = 1.0000 for ternary classification. Discussion and Conclusion: Integrating LLM-derived semantic contextualization into heterogeneous graph learning enables effective modeling of contextual information that is not directly captured by wearable physiological signals alone. The proposed SeGA-GNN framework demonstrates that adaptive semantic fusion substantially improves the accuracy, robustness, and interpretability of wearable-based emotion detection. These findings establish a promising direction for next-generation wearable affective computing systems and intelligent emotion-aware applications.
Colbaugh, R.; Glass, K.
Show abstract
The ubiquity of smartphones in modern life suggests the possibility to use them to continuously monitor patients, for instance to detect undiagnosed diseases or track treatment progress. Such data collection and analysis may be especially beneficial to patients with i.) mental disorders, as these individuals can experience intermittent symptoms and impaired decision-making, which may impede diagnosis and care-seeking, and ii.) progressive neurological diseases, as real-time monitoring could facilitate earlier diagnosis and more effective treatment. This paper presents a new method of leveraging passively-collected smartphone data and machine learning to detect and monitor brain disorders such as depression and Parkinsons disease. Crucially, the algorithm is able learn accurate, interpretable models from small numbers of labeled examples (i.e., smartphone users for whom sensor data has been gathered and disease status has been determined). Predictive modeling is achieved by learning from both real patient data and synthetic patients constructed via adversarial learning. The proposed approach is shown to outperform state-of-the-art techniques in experiments involving disparate brain disorders and multiple patient datasets.
Dong, Z.; Liu, H.; Ge, X.; Zhang, H.; Li, Z.; Chen, Y.; Li, W.
Show abstract
Alzheimers disease (AD) is a progressive neurodegenerative disorder, with mild cognitive impairment (MCI) as its prodromal stage. Accurate MCI conversion prediction is critical for early intervention and resource allocation. Recently, deep learning-based multi-modal neuroimaging fusion has become a hot research topic in AI-assisted AD diagnosis. Existing multimodal fusion approaches are limited by modality heterogeneity, ability to model inter-regional interactions, and insufficient interpretability. To address these challenges, BRC-MMHF, a Brain Region-Centered MultiModal Hypergraph Fusion framework, is proposed. In this framework, parameter-free channel exchange and ROI-level feature extraction mechanisms are employed to reduce modality heterogeneity and extract structurally consistent features from MRI and PET. A multimodal hypergraph models high-order interregional cross-modal relationships, while a lesion-aware module highlights disease-relevant regions to enhance interpretability. Structured clinical data are incorporated through a lightweight tabular encoder to improve adaptability and diagnostic robustness. Experiments on the ADNI dataset show that BRC-MMHF achieves 80.71% accuracy and 89.7% AUC in MCI conversion prediction, outperforming a range of state-of-the-art methods based on MRI and PET imaging, while providing high interpretability.
Li, Z.; Hu, C.; Zhuang, W.; Dong, Z.; Liu, H.; Li, W.
Show abstract
PurposeAccurate disease progression prediction is vital for managing critically ill patients in intensive care. Existing deep learning approaches mainly operate in the time domain and often fail to capture long-range dependencies and spectral dynamics. This study proposes a unified framework integrating time and frequency-domain representations to improve predictive accuracy. MethodsWe introduce FETT (Frequency-Enhanced TCN-Transformer), a dual-domain forecasting framework that combines discrete wavelet transform-based frequency analysis with transformer-based temporal modeling. Based on an iTransformer backbone, FETT introduces three key innovations: (1) a frequency-aware representation module, (2) a dual-TCN architecture that enhances temporal representation through multi-scale feature extraction and global dependency modeling; and (3) a frequency-aware inverse reconstruction module for clinically interpretable time-domain forecasts. ResultsExperiments on the MIMIC-IV Sepsis-3 cohort show that FETT outperforms state-of-the-art baselines, reducing MSE by up to 20.81% and MAE by 13.55% in 24-hour forecasting tasks. Ablation studies confirmed the complementary contributions of the dual-TCN design and frequency-aware modules. ConclusionFETT effectively integrates time- and frequency-domain information to deliver accurate and interpretable disease progression predictions in critical care. By bridging spectral and temporal representations, it enables early detection of patient deterioration and holds strong potential for advancing proactive ICU monitoring and personalized clinical decision-making.
Choi, S.; Gu, G.; Kim, Y.; Lee, S.; Sim, S.-i.; Jang, Y. M.; Kim, H.
Show abstract
Adhesive electrocardiography (ECG) electrodes used in neonatal intensive care units (NICUs) may cause skin injury in premature infants. Although photoplethysmography (PPG)-based ECG reconstruction has been explored, existing studies have mainly focused on adult data and often rely on direct PPG-to-ECG mapping or artificial signal alignment, which may be unsuitable for neonates with highly variable pulse arrival time (PAT). In this study, we propose an alignment-free RoPE-based dual-stream Transformer for reconstructing missing neonatal ECG segments using concurrent PPG signals and bidirectional ECG context. A total of 52,566 10-second ECG-PPG windows were extracted from 159 NICU patients and split at the patient level to prevent data leakage. The model was designed to learn ECG-PPG temporal coupling without forced synchronization by integrating PPG-derived hemodynamic timing information with lead-specific ECG context. Under a 40% random missing condition, the model achieved a Pearson correlation coefficient of 0.96, mean absolute error of 0.04, and root mean square error of 0.07. It also maintained robust performance under 4.0-second continuous block loss and 60% random patch loss, preserving a PCC of at least 0.90. These findings suggest that the proposed framework may serve as a signal imputation module for maintaining ECG monitoring continuity in NICU environments. Prospective validation is required before clinical diagnostic use.
Bhavna, K.; Ghosh, N.; Banerjee, R.; Roy, D.
Show abstract
1Recent technological advancement in Graph Neural Networks (GNNs) have been extensively used to diagnose brain disorders such as autism (ASD), which is associated with deficits in social communication, interaction, and restricted/repetitive behaviors. However, the existing machine-learning/deep-learning (ML/DL) models suffer from low accuracy and explainability due to their internal architecture and feature extraction techniques, which also predominantly focus on node-centric features. As a result, performance is moderate on unseen data due to ignorance of edge-centric features. Here, we argue that meaningful features and information can be extracted by focusing on meta connectivity between large-scale brain networks which is an edge-centric higher order dynamic correlation in time. In the current study, we have proposed a novel explainable and generalized node-edge connectivity-based graph attention neural network(Ex-NEGAT) model to classify ASD subjects from neuro-typicals (TD) on unseen data using a node edge-centric feature set for the first time and predicted their symptom severity scores. We used ABIDE (I and II) dataset with a large sample size (Total no. of samples = 1500). The framework employs meta-connectivity derived from Theory-of-Mind (ToM), Default-mode Network (DMN), Central executive (CEN), and Salience network (SN) that measure the dynamic functional connectivity (dFC) as a flow across morphing connectivity configurations. To generalize the Ex-NEGAT model, we trained the proposed model on ABIDE I(No. of samples =840) and performed testing on the ABIDE II(no. of samples =660) dataset and achieved 88% accuracy with an F1-score of 0.89. Additionally, we identified symptom severity scores for each individual subjects using their meta-connectivity links between relevant brain networks and passing that to Connectome-based Prediction Modelling (CPM) pipeline to identify the specific large-scale brain networks whose edge connectivity contributed positively and negatively to the prediction. Our approach accurately predicted ADOS-Total, ADOS-Social, ADOS-Communication, ADOS-Module, ADOS-STEREO, and FIQ scores.
E, S.; Wang, C.; Rao, T. D.; Kumar, T. S.
Show abstract
Major depressive disorder (MDD) is a common psychiatric disorder that requires reliable and objective assessment for early clinical intervention. Electroencephalography (EEG) is widely used for this purpose because it provides a non-invasive and low-cost measure of brain activity with high temporal resolution. However, EEG-based depression detection remains challenging due to the nonlinear nature of EEG signals, inter-subject variability, and the limited availability of subject-independent evaluation. To address these issues, this paper proposes a hybrid quantum-classical multiscale long short-term memory with parameterized quantum circuit branches (MS-LSTM-PQC) framework for subject-level EEG-based depression detection. The proposed model extracts temporal representations at multiple scales using parallel LSTM branches and incorporates eyes-closed (EC) and eyes-open (EO) condition information through condition-aware feature fusion. To further enhance the learned representations, scale-specific LSTM features are processed using PQC-based quantum branches implemented with TensorFlow Quantum (TFQ), providing an additional nonlinear feature transformation before classification. Experiments were conducted on the Mumtaz EEG depression dataset using EC-only, EO-only, and merged EC+EO conditions with 1-s, 2-s, and 3-s EEG windows. To reduce subject-level data leakage, all experiments were evaluated using 5-fold and 10-fold GroupKFold validation. The best overall accuracies across the evaluated settings were 92.05% and 95.08% under 5-fold and 10-fold GroupKFold validation, respectively. The 2-s merged EC+EO setting provided the most stable performance across validation protocols. In addition, Integrated Gradients (IG)-based explainability analysis showed that frontal and fronto-central channels, especially Fz, showed higher contributions to the model decision. These results suggest that multiscale temporal learning with quantum-enhanced feature transformation can support subject-level EEG-based depression detection under leakage-controlled evaluation.
Jaquenoud, T.; Keene, S.; Shlayan, N.; Federman, A.; Pandey, G.
Show abstract
A growing number of algorithms are being developed to automatically identify disorders or disease biomarkers from digitally recorded audio of patient speech. An important step in these analyses is to identify and isolate the patients speech from that of other speakers or noise that are captured in a recording. However, current algorithms, such as diarization, only label the identified speech segments in terms of non-specific speakers, and do not identify the specific speaker of each segment, e.g., clinician and patient. In this paper, we present a novel algorithm that not only performs diarization on clinical audio, but also identifies the patient among the speakers in the recording and returns an audio file containing only the patients speech. Our algorithm first uses pretrained diarization algorithms to separate the input audio into different tracks according to nonspecific speaker labels. Next, in a novel step not conducted in other diarization tools, the algorithm uses the average loudness (quantified as power) of each audio track to identify the patient, and return the audio track containing only their speech. Using a practical expert-based evaluation methodology and a large dataset of clinical audio recordings, we found that the best implementation of our algorithm achieved near-perfect accuracy on two validation sets. Thus, our algorithm can be used for effectively identifying and isolating patient speech, which can be used in downstream expert and/or data-driven analyses.
Yin, Z.; Zhu, H.
Show abstract
Existing supervised and self-supervised EEG models mainly learn discriminative or reconstructive representations within individual segments, while the transition information between adjacent EEG segments remains underexplored. In this study, we propose a Multimodal self-supervised EEG World Model for wearable seizure detection. Inspired by Le World Model, the proposed method encodes consecutive EEG segments into a shared latent space and predicts the next-segment latent representation from the current-segment representation conditioned on synchronized physiological information from ECG, EMG, and movement (MOV) signals. A learnable query-based fusion module aggregates the auxiliary multimodal representations into a compact physiological condition, while Sketched Isotropic Gaussian Regularization (SIGReg) is applied to stabilize the latent space and prevent representation collapse. After pretraining, only the pretrained EEG encoder is retained and frozen for linear binary probing, enabling EEG-only downstream seizure detection. We evaluated the proposed model on the SeizeIT2 wearable focal epilepsy dataset using a strict patient-wise training, validation, and test split. The proposed Multimodal EEG World Model achieved an AUPRC of 0.3748 , ROC-AUC of 0.8025 , and balanced accuracy of 0.7308 , ranking first on these three metrics among the ablation studies. It also achieved the highest AUPRC, ROC-AUC, balanced accuracy, and F1-score among the evaluated external baselines. These findings demonstrate that synchronized multimodal physiological information can provide useful contextual information for latent EEG transition learning and improve wearable EEG representation learning.
Shehzad, A.; Zhang, D.; Xia, F.; Yu, S.; Abid, S.; Cheng, X.; Zhou, J.
Show abstract
Functional brain networks play an essential role in the diagnosis of brain disorders by enabling the identification of abnormal patterns and connections in brain activities. Previous methods often rely on whole brain functional connectivity approaches to construct these networks using Functional Magnetic Resonance Imaging (fMRI) data. However, these approaches introduce noise and overlook localized disruptions within specific brain subnetworks, leading to potential misdiagnoses. To address this challenging issue, we propose mBrainGT, a modular brain graph transformer model that focuses on modular functional connectivity (mFC) to improve the diagnosis of brain disorders. Compared to existing methods, mBrainGT constructs and analyses functional brain subnetworks individually, reflecting the inherent structure of the brain. It captures both local features within each modular network and their interactions through self-attention and cross-attention mechanisms. It also learns global interactions via adaptive fusion. We validate mBrainGT on three benchmark datasets (ADNI, PPMI, and ABIDE). The results demonstrate that mBrainGT outperforms existing methods in diagnostic accuracy, providing more robust and precise representations of the brain network essential for accurate disease detection. Our study highlights the potential of modular connectivity-based graph learning in the refinement of brain disorder diagnostics, offering a more precise and biologically relevant representation of functional brain networks.
He, K.; Wang, M.
Show abstract
Physical activity (PA) is a critical, modifiable determinant of health. The relationship between PA behavior and health outcomes has been increasingly examined using objectively measured accelerometer data. Accelerometer{square}derived PA features have been consistently associated with a wide range of health outcomes and have shown potential in predictive modeling. However, traditional summary{square}statistic features are limited in their ability to capture the high{square}resolution and dynamic accelerometer data, and the predictive utility of accelerometer data across diverse health outcomes has not been comprehensively evaluated. Here, we present PABformer, a foundation model for accelerometer data pretrained on the UK Biobank to learn behavior{square}level representations of PA. PABformer employs a channel{square}separation strategy to disentangle heterogeneities inherent in accelerometer data and leverages a multi{square}channel Transformer encoder to extract channel{square}specific representations. We finetuned PABformer on benchmark tasks, including demographic attribute inference and Parkinsons disease classification, where it consistently outperformed baseline models across multiple evaluation metrics. Extending beyond benchmarks, we applied the pretrained model to 157 chronic diseases spanning seven organ systems, as well as all{square}cause mortality. PABformer outperformed traditional covariate{square}based models in 48% of prevalent diseases, 54% of incident diseases, and mortality, with particularly marked improvements in three prevalent and nine incident diseases. These findings establish PABformer as a generalizable and scalable foundation model for accelerometer data, enhancing the utility of accelerometer measurements for health outcome assessment and offering the potential to advance disease risk prediction and enable precision health applications.
Li, D.; Liu, W.; Han, H.
Show abstract
Mislabeled learning for high-dimensional data is essentially important in AI health and relevant fields but rarely investigated in machine learning. In this study, we address the challenge by proposing a novel mislabeled learning algorithm for high-dimensional data: psychiatric map diagnosis and applying it to solve a long-time bipolar disorder and schizophrenia misdiagnosis in psychiatry. The proposed algorithm converts each input high-dimensional SNP sample into a corresponding 2D characteristic image called a psychiatric map through feature self-organizing learning. It can automatically detect mislabeled observations and relabel them with the most likely ground truth before reproducible machine learning besides providing informative visualization for mislabeling detection. Our method attains more accurate and reproducible psychiatry diagnoses, besides discovering latent psychiatry subtypes not reported before. It works well for those datasets with a limited number of samples and achieves leading advantages over the deep learning peers. This study also presents new insight into the pathology of psychiatric disorders by constructing the devolution path of psychiatric states via relative entropy analysis that discloses latent internal transfer and devolution road maps between different psychiatric states. To the best of our knowledge, it is the first study to solve mislabeled learning for high-dimensional data and will inspire more future work in this field.
Yang, Z.; Azimi, I.; Jafarlou, S.; Labbaf, S.; Borelli, J.; Dutt, N.; Rahmani, A.
Show abstract
The adverse effects of loneliness on both physical and mental well-being are profound. Although previous research has utilized mobile sensing techniques to detect mental health issues, few studies have utilized state-of-the-art wearable devices to forecast loneliness and comprehend the physiological manifestations of loneliness and its predictive nature. The primary objective of this study is to examine the feasibility of forecasting loneliness by employing wearable devices, such as smart rings and watches, to monitor early physiological indicators of loneliness. Furthermore, smartphones are employed to capture initial behavioral signs of loneliness. To accomplish this, we employed personalized machine learning techniques, leveraging a comprehensive dataset comprising physiological and behavioral information obtained during our study involving the monitoring of college students. Through the development of personalized models, we achieved a notable accuracy of 0.82 and an F-1 score of 0.82 in forecasting loneliness levels seven days in advance. Additionally, the application of Shapley values facilitated model explainability. The wealth of data provided by this study, coupled with the forecasting methodology employed, possesses the potential to augment interventions and facilitate the early identification of loneliness within populations at risk.
Liu, T.; Liu, X.; Bao, Y.; Li, W.; Lin, G. N.
Show abstract
Non-suicidal self-injury (NSSI) among adolescents is a prevalent mental health problem and an important indicator of potential suicide risk. Early objective identification and neural mechanism analysis are therefore crucial for clinical screening and intervention. Traditional assessments mainly rely on self-report scales and clinical interviews, which are vulnerable to subjective bias, clinical experience, and missed diagnosis. Electroencephalography (EEG), with its non-invasive, low-cost, and high-temporal-resolution characteristics, provides a promising physiological basis for identifying NSSI-related neural abnormalities. However, EEG-based intelligent recognition of adolescent NSSI remains limited, and existing studies often emphasize classification performance while lacking systematic neurophysiological interpretation. To address these issues, this study proposes CGA-NSSI, a lightweight deep learning framework for adolescent NSSI recognition. The model integrates a one-dimensional convolutional neural network, bidirectional gated recurrent unit, and multi-head self-attention mechanism to extract local spatiotemporal EEG features, model long-range temporal dependencies, and focus on key pathology-related time segments and channels. A standardized preprocessing pipeline, together with Mixup augmentation and Focal Loss, is further used to alleviate sample imbalance and improve robustness in small clinical EEG datasets. Experiments on a real-world adolescent clinical EEG dataset show that CGA-NSSI can effectively identify NSSI-related EEG patterns under imbalanced sample conditions. Interpretability and functional connectivity analyses further reveal prefrontal-centered cross-regional network reorganization, excessive static functional coupling, reduced dynamic connectivity fluctuations, and increased abnormal state occupancy. These findings suggest that CGA-NSSI not only improves objective NSSI recognition but also provides neurophysiological evidence for understanding adolescent self-injury.
Huang, Z.; Yu, L.; Herbozo Contreras, L. F.; Kavehei, O.
Show abstract
This study presents a novel, computationally efficient training framework demonstrated through bio-signal processing on edge medical devices. The approach integrates conventional full training with an innovative {micro}-Training technique, wherein the encoder and decoder of a compact model remain frozen while only the middle layer is updated. This design is further enhanced by a novel Future-Guided Self-Distillation mechanism that leverages the models anticipated future state in training to boost performance and improve generalization on unseen data, using electrocardiogram (ECG) signals as the primary case study. Additionally, {micro}-Fine-Tuning facilitates ondevice adaptation under resource-constrained conditions. We validate our framework using in-sample data from the Telehealth Network of Minas Gerais (TNMG) and out-of-sample testing on the China Physiological Signal Challenge 2018 (CPSC) datasets. Experimental results demonstrate that our integrated strategy (combining full training, self-distilled {micro}-Training, and {micro}-Fine-Tuning) consistently matches or surpasses conventional methods while significantly improving computational efficiency and mitigating catastrophic forgetting. Deployment on Radxa Zero hardware underscores the approachs practical applicability and scalability. Moreover, a demonstration incorporating the proposed self-distilled {micro}-Training into standard training procedures reveals performance improvements. This highlights the techniques potential for broader applications beyond medical diagnostics and TinyML systems, paving the way for its integration into existing training mechanisms to elevate overall model performance.
Li, M.; Li, X.; Pan, K.; Geva, A.; Yang, D.; Sweet, S. M.; Bonzel, C.-L.; Panickan, V. A.; Xiong, X.; Mandl, K. D.; Cai, T.
Show abstract
The wealth of valuable real-world medical data found within Electronic Health Record (EHR) systems is particularly significant in the field of pediatrics, where conventional clinical studies face notably high barriers. However, constructing accurate knowledge graphs from pediatric EHR data is challenging due to its limited content density compared to EHR data for the general population. Additionally, knowledge graphs built from EHR data primarily covering adult patients may not suit the unique biomedical characteristics of pediatric patients. In this research, we introduce a graph transfer learning approach aimed at constructing precise pediatric knowledge graphs. We present MUlti-source Graph Synthesis (MUGS), an algorithm designed to create embeddings for pediatric EHR codes by leveraging information from three distinct sources: (1) pediatric EHR data, (2) EHR data from the general population, and (3) existing hierarchical medical ontology knowledge shared across different patient populations. We break down these code embeddings into shared and unshared components, facilitating the adaptive and robust capture of varying levels of heterogeneity across different medical sites through meticulous hyperparameter tuning. We assessed the quality of these code embeddings in recognizing established relationships among pediatric codes, as curated from credible online sources, pediatric physicians, or GPT. Furthermore, we developed a web API for visualizing pediatric knowledge graphs generated using MUGS embeddings and devised a phenotyping algorithm to identify patients with characteristics similar to a given profile, with a specific focus on pediatric pulmonary hypertension (PH). The MUGS-generated embeddings demonstrated resilience against negative transfer and exhibited superior performance across all three tasks when compared to pediatric-only approaches, multi-site pooling, and semantic-based methods. MUGS embeddings open up new avenues for evidence-based pediatric research utilizing EHR data.
Zhou, M.; Zhang, M.; Wang, J.; Shao, C.; Yan, G.
Show abstract
Cardiovascular disease is one of the leading causes of death worldwide, with myocardial infarction (MI) being a major cause of both morbidity and mortality among cardiovascular patients. MI Patients face a higher risk of cardiovascular disease recurrence afterwards. Therefore, accurately predicting the risk of recurrence and identifying key risk factors are crucial for clinical decision-making. In this paper, we consider the interrelationships among cardiovascular factors from a systemic perspective. We first construct a differential network for each patient to capture individual-specific deviations in factor relationships and propose a novel method, termed Causal Factor-aware Graph Neural Network (CFGNN), which integrates factor interactions to predict the recurrence risk of MI patients while uncovering key risk factors from a causal perspective. Experimental results demonstrate that CFGNN performs well on hospital-derived datasets in real world, effectively identifying several key risk factors. This method not only deepens our understanding of cardiovascular disease, but also paves the way for more targeted and effective interventions.
Akhila, N.; Ekbal, A.; Roy, D.
Show abstract
Accurate diagnosis of Parkinson's disease (PD) remains challenging due to substantial inter-subject variability and the absence of widely accessible, objective multimodal biomarkers. Although speech and magnetoencephalography (MEG) biomarkers have individually demonstrated strong discriminative potential, their joint utilization is constrained by the absence of subject-level paired datasets - a fundamental gap that has prevented cross-modal validation at the individual level. We argue that this makes cross-cohort representation learning not merely a pragmatic workaround, but the most realistic and clinically transferable framework for multimodal PD assessment. In real-world deployment, acoustic screening and neuroimaging biomarkers are acquired through separate clinical pathways and must be integrated across heterogeneous patient populations. To address this, we propose MIRA-Net (Modality-Invariant Residual Adversarial Network). This cross-cohort representation learning framework integrates acoustic speech features from four established UCI datasets (n = 193) with beta-band MEG biomarkers from the NatMEG-PD dataset (n = 127) for PD classification. MIRA-Net employs RF-SHAP feature selection, gradient-reversal-based domain adaptation, and supervised contrastive alignment to learn participant-independent, modality-invariant embeddings. The framework is evaluated under Rest, Go, and Passive task conditions against Early Fusion, Vanilla DANN, and Supervised Contrastive Learning baselines. MIRA-Net achieves a peak accuracy of 86.23% (Go condition, Stacking classifier) with AUC values exceeding 0.88 under repeated cross-validation, alongside a sensitivity of 89.4% and specificity of 83.1%. Friedman tests confirm statistically significant performance differences among fusion strategies (p < 0.003 across all conditions). These results demonstrate that cross-cohort representation learning can extract robust disease-discriminative signatures without synchronized multimodal recordings, offering a practical pathway toward AI-assisted PD assessment in resource-constrained clinical settings.