Back

IEEE Access

Institute of Electrical and Electronics Engineers (IEEE)

All preprints, ranked by how well they match IEEE Access's content profile, based on 35 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Head and Body Pose Classification for Understanding Sleep Behaviour in People Living with Dementia using Video and a Novel Multi-Head Attention-Driven Deep Learning Architecture

Al-Gawwam, S.; M Pineda, M.; K G Ravindran, K.; della Monica, C.; Atzori, G.; Nilforooshan, R.; Hassanin, H.; Revell, V.; Dijk, D.-J.; Wells, K.

2026-05-06 health economics 10.64898/2026.04.29.26351379 medRxiv
Top 0.1%
54.8%
Show abstract

Sleep posture is known to be relevant to various sleep disorders, such as sleep apnea, but is not often quantified in sleep monitoring systems. We address this with a novel vision-based approach, which is robust to the challenging conditions (variable lighting, partial occlusions, variable geometry) of inbed monitoring. This paper proposes a novel, attention-driven deep learning framework for the robust classification of head and body pose from infrared (IR) video streams during sleep of older people and people living with Alzheimers. Our architecture integrates a pre-trained convolutional backbone with a novel Multi-Head Channel-Spatial Attention (MH-CSA) module. The MH-CSA mechanism hierarchically identifies salient features by first capturing multi-scale spatial context using parallel heads with varied dilation rates, and then adaptively recalibrating feature importance via integrated Squeeze-and-Excitation blocks. To specifically address class imbalance, the model is optimized using a Dynamic Class-Balanced Focal Loss, which forces the network to focus on hard-to-classify examples from underrepresented classes. Whilst most prior sleep analysis work is developed using data from healthy younger participants, our system was developed and validated on a nocturnal sleep dataset of older adults and people living with Alzheimers disease, with IR video synchronized to clinical video-Polysomnography (vPSG). For head position classification, the system achieved an F1-score of 91% for older adults and 90% for people living with Alzheimers; for body pose prediction, the scores were 91% and 89% for the respective cohorts. These results demonstrate significant potential for application in understanding sleep behavior and informing appropriate sleep interventions.

2
Patient Clustering and Classification for Vital Organ Failure Using ICD Code with Graph Attention

Liu, Z.; Hu, Y.; Mertes, G.; Yang, Y.; Clifton, D.

2022-11-07 bioengineering 10.1101/2022.11.07.515209 medRxiv
Top 0.1%
32.8%
Show abstract

ObjectiveHeart failure, respiratory failure and kidney failure are three severe organ failures (OF) that have high mortalities and are most prevalent in intensive care units. The objective of this work is to offer insights on OF clustering from the aspects of graph neural network and diagnosis history. MethodsThis paper proposes a neural network-based pipeline to cluster three types of organ failure patients by incorporating embedding pre-train using an ontology graph of International Classification of Diseases (ICD) codes. We employ an autoencoder-based deep clustering architecture jointly trained with a K-means loss, and a non-linear dimension reduction is performed to obtain patient clusters on the MIMIC-III dataset. ResultsThe clustering pipeline shows superior performance on a public-domain image dataset. For MIMIC-III, the model gives two distinct clusters that are related to the severity of the diseases. The learnt ICD embeddings present strong power in identifying the OF type in supervised learning. ConclusionOur proposed pipeline gives stable clusters, however, they do not correspond to the type of OF which indicates these OF share significant hidden characteristics in diagnosis. These clusters can be used to signal possible complications and severity of illness. SignificanceWe are the first to apply an unsupervised approach to offer insights from a biomedical engineering perspective on these three types of organ failure, and publish the pre-trained embeddings for future transfer learning.

3
Use of Machine Learning for Long Term Planning and Cost Minimization in Healthcare Management

kabir, s. b.; Shuvo, S. S.; Ahmed, H. U.

2021-10-07 health economics 10.1101/2021.10.06.21264654 medRxiv
Top 0.1%
31.6%
Show abstract

The Healthcare system of a country is a crucial infrastructure that requires long-term capacity planning. The covid 19 outbreak pointed to the necessity of adequate hospital capacity, especially for developing countries like Bangladesh. The existing infrastructure planning of these countries emphasizes short-term goals and lacks vision planning for a long time horizon. It is in the countrys best interest to make long-term capacity expansion plans, a strategy the developed countries banked to provide adequate healthcare facilities to their residents. However, no single solution is appropriate for a different region. Hence, it is required to comprehensively study the situation and constraints of the specific region before providing expensive capacity expansion plans. This work focuses on applying a deep Reinforcement Learning based long-term hospital bed capacity expansion plan. We utilize the RNN-LSTM based population forecast, deep Reinforcement Learning (RL) based policy-making, and state-of-the-art Artificial Intelligence techniques to provide a solution. We perform a case study for the Abhaynagar Upazila of Jessore, one of the largest cities in the southwest part of Bangladesh, to analyze the benefits of such an approach compared to existing myopic policies. The experiment results show that the deep RL-based policy significantly minimizes cost over a 30-year expansion plan.

4
Hand Drawing Image based Causal Representation Learning for Robust Parkinson's Disease Feature Extraction and Detection

Kim, D.

2025-08-14 bioengineering 10.1101/2025.06.01.657220 medRxiv
Top 0.1%
31.3%
Show abstract

Being an irreversible disorder regarding the human motor-system, Parkinsons Disease(PD) has been a threat to many neurological patients, especially due to its severity in pain and muscle control restriction. As PD has no significant cure or treatments to this day, early diagnosis, or detections of PD within potential patients is a crucial task to maximize the effect of mediations which are implemented to achieve temporal prohibition of motor failure progression. In recent research, alongside conventional diagnosis methods based on neurological examinations or MRI based brain imaging, use of deep learning based artificial intelligence models, such as ResNet, are repeatedly reported to have significant progress in detecting PD in early stages with high performance. Based on current success, this research attempted to further enhance AI-driven PD diagnosis by developing a deep learning based causal representation learning framework that extracts only highly robust PD features from simple hand drawings. Specifically, convolutional VAE based reconstruction and information theory based weakly supervised learning were linked with causal representation learning methods to distinguish significant PD features from geometrical features within hand drawing images. Not only aiding conventional tests for PD diagnosis, but also giving reliable representations of PD features such as tremor and rigidity, developed framework was found to achieve high performance in both retrieving latent factors of PD in images and predicting PD diagnosis results.

5
An automatic and efficient pulmonary nodule detection system based on multi-model ensemble

Chen, J.; Wang, W.; Ju, B.; Jiang, J.; Zhang, L.; He, H.; Zhang, X.; Shen, Y.

2020-04-20 bioengineering 10.1101/2020.04.14.040931 medRxiv
Top 0.1%
28.4%
Show abstract

Accurate pulmonary nodule detection plays an important role in early screening of lung cancer. Although there are many presented CAD systems based on deep learning for pulmonary nodule detection, these methods still have some problems in clinical use. The improvement of false negatives rate of tiny nodules, the reduction of false alarms and the optimization of time consumption are some of them that need to be solved as soon as possible. In view of the above problems, in this paper, we first propose a novel full convolution segmentation framework for lung cavity extraction in preprocessing stage to solve the time consumption problem of the existing pulmonary nodule detection systems. Furthermore, a 2D-NestedUNet segmentation network and a 3D-RPN detection network is stacked to get the high recall and low false positive rate on nodule candidate extraction, especially the recall of tiny nodules. Finally, a false positive reduction method based on multi-model ensemble is proposed for the further classification of nodule candidates. Our methods are evaluated on several public datasets, LUNA16, LNDb and ChestCT2019, which demonstrated the superior performance of our CAD system.

6
XAI-based Data Visualization in Multimodal Medical Data

Sharma, S.; Singh, M.; McDaid, L.; Bhattacharyya, S.

2025-07-15 bioengineering Community evaluation 10.1101/2025.07.11.664302 medRxiv
Top 0.1%
26.6%
Show abstract

Explainable Artificial Intelligence (XAI) is crucial in healthcare as it helps make intricate machine learning models understandable and clear, especially when working with diverse medical data, enhancing trust, improving diagnostic accuracy, and facilitating better patient outcomes. This paper thoroughly examines the most advanced XAI techniques used in multimodal medical datasets. These strategies include perturbation-based methods, concept-based explanations, and example-based explanations. The value of perturbation-based approaches such as LIME and SHAP in explaining model predictions in medical diagnostics is explored. The paper discusses using concept-based explanations to connect machine learning results with concepts humans can understand. This helps to improve the interpretability of models that handle different types of data, including electronic health records (EHRs), behavioural, omics, sensors, and imaging data. Example-based strategies, such as prototypes and counterfactual explanations, are emphasised for offering intuitive and accessible explanations for healthcare judgments. The paper also explores the difficulties encountered in this field, which include managing data with high dimensions, balancing the tradeoff between accuracy and interpretability, and dealing with limited data by generating synthetic data. Recommendations in future studies focus on improving the practicality and dependability of XAI in clinical settings.

7
Beta regression with spatio-temporal effects as a tool for hospital impact analysis of initial phase epidemics: the case of COVID-19 in Spain

Jannes, G.; Barreal, J.

2020-06-29 health economics 10.1101/2020.06.27.20141614 medRxiv
Top 0.1%
23.4%
Show abstract

COVID-19 has put an extraordinary strain on medical staff around the world, but also on hospital facilities and the global capacity of national healthcare systems. In this paper, Beta regression is introduced as a tool to analyze the rate of hospitalization and the proportion of Intensive Care Unit admissions over both hospitalized and diagnosed patients, with the aim of explaining as well as predicting, and thus allowing to better anticipate, the impact on hospital resources during an early-phase epidemic. This is applied to the initial phase COVID-19 pandemic in Spain and its different regions from 20-Feb to 08-Apr of 2020. Spatial and temporal factors are included in the Beta distribution through a precision factor. The model reveals the importance of the lagged data of hospital occupation, as well as the rate of recovered patients. Excellent agreement is found for next-day predictions, while even for multiple-day predictions (up to 12 days), robust results are obtained in most cases in spite of the limited reliability and consistency of the data.

8
A CNN-LSTM Architecture for Detection of Intracranial Hemorrhage on CT Scans

Nguyen, T. N.; Tran, Q. D.; Nguyen, T. N.; Nguyen, Q. H.

2020-04-22 health economics 10.1101/2020.04.17.20070193 medRxiv
Top 0.1%
23.1%
Show abstract

We propose a novel method that combines a convolutional neural network (CNN) with a long short-term memory (LSTM) mechanism for accurate prediction of intracranial hemorrhage on computed tomography (CT) scans. The CNN plays the role of a slice-wise feature extractor while the LSTM is responsible for linking the features across slices. The whole architecture is trained end-to-end with input being an RGB-like image formed by stacking 3 different viewing windows of a single slice. We validate the method on the recent RSNA Intracranial Hemorrhage Detection challenge and on the CQ500 dataset. For the RSNA challenge, our best single model achieves a weighted log loss of 0.0522 on the leaderboard, which is comparable to the top 3% performances, almost all of which make use of ensemble learning. Importantly, our method generalizes very well: the model trained on the RSNA dataset significantly outperforms the 2D model, which does not take into account the relationship between slices, on CQ500. Our codes and models will be made public.

9
Heterogeneity analysis of acute exacerbations of chronic obstructive pulmonary disease and a deep learning framework with weak supervision and privacy protection

Suzuki, Y.; Hill, A.; Engel, E.; Lyng, G.; Granchelli, A.; Lockhart, G.; Banaei-Kashani, F.; Bowler, R. P.

2023-12-07 bioinformatics 10.1101/2023.12.04.570028 medRxiv
Top 0.1%
22.6%
Show abstract

11.1 BackgroundChronic obstructive pulmonary disease (COPD) affects 5-10% of the adult US population and is a major cause of mortality. Acute exacerbations of COPD (AECOPDs) are a major driver of COPD morbidity and mortality, but there are no cost-effective methods to identify early AECOPDs when treatment is most likely to reduce the severity and duration of AECOPDs. 1.2 MethodsWe conducted the first long-term (> 12 months), real-time monitoring studies of AECOPD with wearable sensors and self-reporting. We applied a deep learning-based autoencoder for feature extractions, then applied K-means clustering to detect heterogeneity. Accordingly, we proposed a weakly supervised active learning framework to develop anomaly detection models for robust identification of early AECOPD, and a clustered federated learning approach to personalize the anomaly detection models for early detection of heterogeneous subtypes of AECOPD. We evaluated this model by comparing it with other unsupervised learning models and federated learning models. 1.3 FindingsWe identified two clusters based on the Silhouette score and SHAP analysis.One cluster shows high heart rate, low calories, and low steps; and the other has opposite characteristics. We also found out that a single subject could have exacerbation events from both clusters, indicating that there is not only subject-level heterogeneity but also event-level heterogeneity. Our weakly supervised framework outperformed unsupervised methods by 0.06 in average precision with 25 human annotation labels per subject. Our federated learning framework outperformed standard federated learning methods by 0.14 in F1 score and 0.17 in average precision. 1.4 InterpretationWe showed subject-level and event-level heterogeneity in AECOPD using mobile and wearable device data and developed a practical AECOPD detection framework with limited human annotated labels and keeping data private in each device.

10
Hierarchical Coarse-to-Fine cGAN for Subtype-Specific Freezing of Gait Signal Generation

Yu, X.; Cockx, H.; Wang, Y.; Wezel, R. v.; Martens, K. E.; Arami, A.

2025-10-04 bioengineering 10.1101/2025.10.04.680444 medRxiv
Top 0.1%
22.3%
Show abstract

Freezing of gait (FOG), a debilitating symptom of Parkinsons disease, manifests in subtypes as shuffling, trembling, or akinesias, with occurrence and frequency varying across patients. While deep learning (DL) models show promise in FOG detection, their robustness and generalization across subtypes are limited by data scarcity and imbalances between FOG/non-FOG classes and among subtypes. To address this, we propose a subtype-aware FOG augmentation technique enabling training of DL models to perform consistently across subtypes. Specifically, we introduce Hierarchical Coarse-to-Fine conditional Generative Adversary Network (Hi-CF cGAN), a two-stage model that generates subtype-conditioned FOG-like ankle accelerations that are realistic and diverse, as verified through visualization, UMAPs, and Maximum Mean Discrepancy comparison against real signals. We evaluate its effectiveness by training CNNs for FOG detection with both general (subtype-stratified) and personalized (subtype-variant, based on patient-specific subtype composition) augmentation via Hi-CF cGAN, benchmarking against classical augmentations and baseline (no augmentation). Compared to baseline, general augmentation with Hi-CF cGAN effectively improves average detection rates of FOG, trembling FOG, and especially the previously overlooked minor subtypes, shuffling FOG (from 66.8% to 81.6%) and akinesia FOG (from 58.7% to 77.9%). These improvements exceed those of classical augmentations, demonstrating superior realism, richness, and adaptability of Hi-CF cGAN-generated data in addressing FOG/non-FOG and subtype imbalances. Personalized augmentation further enhances accuracy on targeted subtype(s) compared to general augmentation, highlighting its potential for tailored model optimization.

11
Graph Autoencoder and StrNN based Causal Analysis of Mortality in Heart Failure Patients

Kim, D.

2025-04-23 bioengineering 10.1101/2024.11.11.622921 medRxiv
Top 0.1%
19.9%
Show abstract

Though analyzed for decades, dissecting and finding mechanisms of cardiovascular diseases, especially heart failures, are still an on-going task for many researchers. However, through recent floods of machine learning and deep learning algorithms to replace traditional approaches, and their applications in diverse cardiovascular research areas, it seems plausible to say that conquering or preventing heart failure catastrophes might no longer be a delusional task within a few more years. To accelerate the arrival of a new era, this research implemented several cutting-edge algorithms currently introduced in causal deep learning to observational heart disease patient data to find key mechanisms that lead to cardiac deaths under a highly flexible framework. Extracting latent causal DAGs from observational data using Graph Auto Encoder, and finding specific causal relationships and interventional effects under Structured Neural Networks (StrNN), novel findings regarding key causes of deaths in heart failure patients were found in numerous aspects. Specifically, existence of intervals where average treatment effects due to causal interventions in platelets, ejection fraction, and serum creatinine levels dramatically decrease or increase was found among heart patients, which can lead to significant eliminations or additions of practical clinical treatments in terms of reducing cardiac death event probability after cardiac failure.

12
COVID-19 Detection on Chest X-Ray and CT Scan Images Using Multi-image Augmented Deep Learning Model

Purohit, K.; Kesarwani, A.; Kisku, D. R.; Dalui, M.

2020-10-19 bioengineering 10.1101/2020.07.15.205567 medRxiv
Top 0.1%
19.8%
Show abstract

COVID-19 is posed as very infectious and deadly pneumonia type disease until recent time. Despite having lengthy testing time, RT-PCR is a proven testing methodology to detect coronavirus infection. Sometimes, it might give more false positive and false negative results than the desired rates. Therefore, to assist the traditional RT-PCR methodology for accurate clinical diagnosis, COVID-19 screening can be adopted with X-Ray and CT scan images of lung of an individual. This image based diagnosis will bring radical change in detecting coronavirus infection in human body with ease and having zero or near to zero false positives and false negatives rates. This paper reports a convolutional neural network (CNN) based multi-image augmentation technique for detecting COVID-19 in chest X-Ray and chest CT scan images of coronavirus suspected individuals. Multi-image augmentation makes use of discontinuity information obtained in the filtered images for increasing the number of effective examples for training the CNN model. With this approach, the proposed model exhibits higher classification accuracy around 95.38% and 98.97% for CT scan and X-Ray images respectively. CT scan images with multi-image augmentation achieves sensitivity of 94.78% and specificity of 95.98%, whereas X-Ray images with multi-image augmentation achieves sensitivity of 99.07% and specificity of 98.88%. Evaluation has been done on publicly available databases containing both chest X-Ray and CT scan images and the experimental results are also compared with ResNet-50 and VGG-16 models.

13
Pneumonia Detection with Semantic Similarity Scores

Gholamipoor, r.; Rafiee, N.; Kollmann, M.

2021-10-15 bioinformatics 10.1101/2021.10.14.464247 medRxiv
Top 0.1%
19.1%
Show abstract

X-ray images have been widely used for medical diagnoses of cardiothoracic and pulmonary abnormalities due to its noninvasiveness. Advancement in computer-aided diagnostic technologies, such as deep supervised methods, can help radiologists with a reliable early treatment and reduce diagnosis time. Nevertheless, these methods are prone to the small number of labeled samples and are limited to a specific abnormality. In this paper we combined a self-supervised contrastive method with a Mahalanobis distance score to develope an abnormality detection method that uses only healthy images during the training procedure. We were able to outperform previous unsupervised methods for the task of Pneumonia detection. We show that representation learned by the self-supervised method improves the supervised tasks for Pneumonia detection.

14
How to Make COVID-19 Contact Tracing Apps work: Insights From Behavioral Economics

Ayres, I.; Romano, A.; Sotis, C.

2020-09-11 health economics 10.1101/2020.09.09.20191320 medRxiv
Top 0.1%
19.1%
Show abstract

Due to network effects, Contact Tracing Apps (CTAs) are only effective if many people download them. However, the response to CTAs has been tepid. For example, in France less than 2 million people (roughly 3% of the population) downloaded the CTA. Against this background, we carry out an online experiment to show that CTAs can still play a key role in containing the spread of COVID-19, provided that they are re-conceptualized to account for insights from behavioral science. We start by showing that carefully devised in-app notifications are effective in inducing prudent behavior like wearing a mask or staying home. In particular, people that are notified that they are taking too much risk and could become a superspreader engage in more prudent behavior. Building on this result, we suggest that CTAs should be re-framed as Behavioral Feedback Apps (BFAs). The main function of BFAs would be providing users with information on how to minimize the risk of contracting COVID-19, like how crowded a store is likely to be. Moreover, the BFA could have a rating system that allows users to flag stores that do not respect safety norms like wearing masks. These functions can inform the behavior of app users, thus playing a key role in containing the spread of the virus even if a small percentage of people download the BFA. While effective contact tracing is impossible when only 3% of the population downloads the app, less risk taking by small portions of the population can produce large benefits. BFAs can be programmed so that users can also activate a tracing function akin to the one currently carried out by CTAs. Making contact tracing an ancillary, opt-in function might facilitate a wider acceptance of BFAs.

15
Labelling Human Kinematics Data Using Classification Models

Shi, Y.; Chadderwala, N.; Ratan, U.

2022-02-24 rehabilitation medicine and physical therapy 10.1101/2022.02.18.22271206 medRxiv
Top 0.1%
19.0%
Show abstract

The goal of this study is to develop a classification model that can accurately and efficiently label human kinematics data. Kinematics data provides information about the movement of individuals by placing sensors on the human body and tracking their velocity, acceleration and position in three dimensions. These data points are available in C3D format that contains numerical data transformed from 3D data captured from the sensors. The data points can be used to analyse movements of injured patients or patients with physical disorders. To get an accurate view of the movements, the datasets generated by the sensors need to be properly labelled. Due to inconsistencies in the data capture process, there are instances where the markers have missing data or missing labels. The missing labels are a hindrance in motion analysis as it introduces noise and produces incomplete datapoints of sensors positioning in 3 dimensional space. Labelling the data manually introduces substantial effort in the analysis process. In this paper, we will describe approaches to pre-process the kinematics data from its raw format and label the data points with missing markers using classification models.

16
Semantic-Aware Energy-Efficient Operation inSmart Capsule Endoscopy

Zoofaghari, M.; Rahaimifard, A.; Chatterjee, S.; Balasingham, I.

2026-03-19 bioinformatics 10.64898/2026.03.17.712375 medRxiv
Top 0.1%
18.9%
Show abstract

Goal-oriented semantic communication has recently emerged in wireless sensor-actuator networks, emphasizing the meaning and relevance of information over raw data delivery, thereby enabling resource-efficient telecommunication. This paradigm offers significant benefits for intra-body or implantable sensor-actuator networks, including dramatic reductions in bandwidth requirements, latency, and power consumption. In this paper, we address a patch-based energy-efficient anomaly detection method for smart capsule endoscopy. We propose a deep learningbased algorithm that employs the similarity between features extracted from measured images and a reference (normal) image as the detection metric. The algorithm is evaluated using a clinical dataset of capsule-captured images, combined with a simulated intra-body channel model. The results demonstrate that even with only 60% of the transmission power (relative to a standard link design for QPSK modulation) and 65% of the light intensity, the probability of anomaly detection remains above 85%, and it gradually improves as power and illumination levels increase. This improvement translates into a potential battery life extension of over 43%. The findings highlight the potential of semanticaware, energy-efficient intra-body devices for more sustainable and effective medical interventions.

17
A Neural Network for High-Precise and Well-Interpretable Electrocardiogram Classification

Liu, X.; Liu, X.

2024-01-04 bioengineering 10.1101/2024.01.03.573822 medRxiv
Top 0.1%
18.9%
Show abstract

Manual heart disease diagnosis with the electrocardiogram (ECG) is intractable due to the intertwined signal features and lengthy diagnosis procedure, especially for the 24-hour dynamic ECG signals. Consequently, even experienced cardiologists may face difficulty in producing all accurate ECG reports. In recent years, neural network-based automatic ECG diagnosis methods have exhibited promising performance, suggesting a potential alternative to the labor-intensive examination conducted by cardiologists. However, many existing approaches failed to adequately consider the temporal and channel dimensions when assembling features and ignored interpretability. And clinical theory underscores the necessity of prolonged signal observations for diagnosing certain ECG conditions such as tachycardia. Moreover, specific heart diseases manifest primarily through distinct ECG leads represented as channels. In response to these challenges, this paper introduces a novel neural network architecture for ECG classification (diagnosis). The proposed model incorporates Lead Fusing blocks, transformer-XL encoder-based Encoder modules, and hierarchical temporal attentions. Importantly, this classifier operates directly on raw ECG time-series signals rather than cardiac cycles. Signal integration begins with the Lead Fusing blocks, followed by the Encoder modules and hierarchical temporal attentions, enabling the extraction of long-dependent features. Furthermore, we argue that existing convolution-based methods compromise interpretability, while our proposed neural network offers improved clarity in this regard. Experimental evaluation on a comprehensive public dataset confirms the superiority of our classifier over state-of-the-art methods. Moreover, visualizations reveal the enhanced interpretability provided by our approach. HighlightsO_LIOur model extracts long-dependent features of ECG signals based on the Transformer-XL encoder. C_LIO_LIThe proposed network offers the improved interpretability. C_LIO_LIOur classifier achieves superior performance over other state-of-the-art methods. C_LI

18
GuavaVision AI: An Explainable Deep Learning Framework for Automated Classification, Lesion Localization, and Segmentation of Guava Diseases

Biswas, J.; Islam, M.; Bangabashi, M. M.; Akter, M.; Nishi, T. S.; Sheikh, M. K.; Mia, M. R.; Anwar, M. M.

2026-06-23 bioengineering 10.64898/2026.06.18.733093 medRxiv
Top 0.1%
18.8%
Show abstract

Guava cultivation is considerably influenced by foliar and fruit diseases whose overlapping symptoms and environmental variability make accurate field-level diagnosis challenging. Numerous studies have been conducted to find efficient methods of diagnosing plant diseases, but most focus on image-level classification and do not include lesion localization or pixel-level segmentation of the images within a single framework of analysis. This study proposes a comprehensive framework for utilizing automated image analysis to classify guava leaf and fruit diseases at the image level, locate lesions, and segment lesions at the pixel level from multiple images of the same type of disease collected from various growing conditions. The dataset was enriched through three augmentation strategies including standard preprocessing, structured augmentation, and GAN-based synthetic image generation, expanding the effective training data to approximately 7,000 images, while a 5-fold cross-validation strategy guided model selection and final performance was assessed on a held-out test set. The experimental evaluation of multiple state-of-the-art Convolutional Neural Networks (CNNs) for the classification of guava leaf and fruit diseases indicated that the model generated using the ResNet50+DenseNet121 model fusion achieved the highest classification accuracy of 98.20%. For lesion detection and segmentation, YOLOv8-seg outperformed Mask R-CNN, achieving mAP@0.5 of 0.907 and 0.889, and mAP@0.5:0.95 of 0.783 and 0.769 for detection and segmentation, respectively, with a balanced precision-recall profile. The techniques of Explainable AI (XAI) were used to increase the transparency of this model by identifying areas in the image that are significant to the actual lesion. The framework was further designed with practical web-based deployment in mind, evaluating both lightweight and high-capacity models to balance computational efficiency against predictive accuracy. From this research, it was concluded that using model fusion, data augmentation, and segmentation-aware lesion detection would provide a solution for managing guava diseases effectively.

19
CVAE-based Causal Representation Learning from Retinal Fundus Images for Age Related Macular Degeneration(AMD) Prediction

Kim, D.

2025-02-17 bioengineering 10.1101/2025.02.13.638092 medRxiv
Top 0.1%
18.7%
Show abstract

Regarding relatively poor prognosis and acute vision impairment, analyzing Age-Related Macular Degeneration, or AMD has been one of the most important tasks in retinal disease analysis. Especially, constructing methods to analyze and predict Wet AMD, which is characterized by rapid RPE damage due to neovascularization, has been a demanding task for many ophthalmologists for decades. Recently, with advancements in ML/DL frameworks and computer vision AI, these previous efforts are now leading to drastic enhancements in AMD prediction and mechanism analysis. Specifically, use of attention mechanism based CNNs or XAI methods are leading to higher performance in predicting AMD status and reliable explanations. Under current success in the usage of cutting-edge techniques, this research implemented a novel latent causal representation learning framework to further enhance AI-based models to comprehend complex causal AMD mechanisms with only access to retinal fundus images, while constructing a more reliable type of AMD prediction model. Results show that valid convolutional VAE and GAE based explicit latent causal modeling can lead to successful causal disentanglements of underlying AMD mechanisms, while returning essential causal factors that can be utilized to reliably distinguish normal fundus and AMD fundus images in downstream tasks such as diagnosis prediction.

20
Measures of Behavior and Life Dynamics from Commonly Available GPS Data (DPLocate): Algorithm Development and Validation

Rahimi Eichi, H.; Coombs, G.; Baker, J. T.; Onnela, J.-P.; Buckner, R. L.

2022-07-10 psychiatry and clinical psychology 10.1101/2022.07.05.22277276 medRxiv
Top 0.1%
18.6%
Show abstract

Locations of people moving about their lives are now commonly tracked through smartphones and wearable devices that access the Global Positioning System (GPS). Immediate measures include the estimated locations that identify visited map points and the travel paths between them. Here we introduce DPLocate, an open-source GPS data analysis pipeline designed to derive measures that abstract away from the original locations (and hence the identity of the individuals) and capture dynamics related to social, vocational, sleep, and clinical behaviors. We divide derived measures into primary and secondary. Primarily derived measures stay close to the original location data and extract deidentified metrics, including distance traveled, time spent at the main locations, and estimates of travel activity (entropy). Secondary derived measures estimate life patterns that are captured incidentally by extracting returns to the Points of Interest (POIs) in behaviorally-relevant time-bands. For example, measures of behavioral dynamics and social interactions can be gleaned by estimating the time spent in POIs across day, evening, night, and late-night time-bands. The utility of these derived measures for research is illustrated in college students and for clinical monitoring in individuals living with psychiatric disorders. Captured dynamics included behavioral transitions at the onset of the Covid-19 lockdown. Limitations of derived data are discussed, including the necessity to protect derived data from identification and possible ways in which the derived data might be misinterpreted.