Back

Ibis

Wiley

Preprints posted in the last 7 days, ranked by how well they match Ibis's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Southern (California) Sea Otter Population Status and Trends at San Nicolas Island, 2023-Winter 2026

Tomoleoni, J. A.; Yee, J. L.; Seacord, E.; Staedler, M. M.; Hatfield, B. B.; Carswell, L.; Fujii, J.; Bentall, G. B.; Konrad, L.; Young, C.; Tinker, M. T.; Bowen, L.

2026-08-31 ecology 10.64898/2026.08.28.747881 medRxiv
Top 0.6%
0.2%
Show abstract

The southern sea otter (Enhydra lutris nereis) population at San Nicolas Island, California, has been monitored annually since the translocation of 140 sea otters to the island was completed in 1990. Monitoring efforts have varied in frequency and method across years. In 2017, in accordance with the National Defense Authorization Act for Fiscal Year 2016, the U.S. Navy and the U.S. Fish and Wildlife Service formally initiated a sea otter monitoring and research plan to determine the effects of military readiness activities on the growth or decline of the southern sea otter population at San Nicolas Island. The monitoring program, at its basic level, includes quarterly seasonal surveys of population abundance, distribution, and foraging activity. This report presents data from the program with a focus on the recent three years from winter 2023 through winter (February) 2026. From 2023 to 2026, we measured an 8.1-percent per annum decrease in population abundance (95-percent confidence interval =1.1-14.6 percent), with 106 total individuals counted as of February 2026. Historically, sea otter habitat usage at San Nicolas Island was concentrated on the west end of the island. Between 2017 and 2019, we observed increased seasonal usage of the north and south sides of the island, and in 2020-2022, a large (approximately 30-40 individuals) group of sea otters (raft) took up residence off the east end. During 2023-2026 the east end raft disappeared, and sea otters returned to their historical habitat usage patterns at the west end of the island. Foraging data were collected from summer 2023 to winter 2026 on a total of 461 foraging dives in 32 foraging bouts, and the majority of identified prey on successful dives (n=325) were sea urchins (124) followed by snails (48), bivalves (41) and crabs (23). One lobster and one octopus were also identified among the sea otter prey items. We combined these data with data from 2020-2022 to estimate overall energy intake rates that averaged 7.7 kilocalories per minute (95-percent credible interval =6.6-9.1 kilocalories per minute). These results can be useful to the planning of future monitoring and research of sea otters at San Nicolas Island.

2
Copulation calls indicate fertility but do not reflect female mate competition in wild Guinea baboons

Niederbremer, C.; Dal Pesco, F.; Mundry, R.; Neumann, C.; Diakhate, N.; Fischer, J.

2026-09-01 animal behavior and cognition 10.64898/2026.08.26.747217 medRxiv
Top 0.8%
0.2%
Show abstract

Across different modalities, signals play a core role in attracting mates and influencing mating success. In several non-human primate species, females produce calls during mating that are thought to promote male competition over receptive females. The extent to which social system characteristics modulate the function of copulation calls remains less clear. We studied copulation calls in wild Guinea baboons (Papio papio), who live in a multilevel society structured around units in which females associate and mate almost exclusively with a single male. We hypothesised that females use copulation calls as an indirect form of mate competition, with competition increasing in larger units. In addition, we hypothesised that females are more likely to mate again after calling. We analysed 6116 copulations between 2014 and 2025, involving 99 reproductively active females and 78 subadult and adult males. Females produced copulation calls in 72.7% of copulations, with large inter-individual variation. Neither unit size nor its interaction with the female's swelling size or the presence of simultaneously receptive females affected the probability of calling. A survival analysis with a subset of the data (2353 copulations) revealed no effect of calling on the latency to the next mating. Our results render the hypothesis that female Guinea baboons use calls in indirect mate competition unlikely. Yet, the probability of calling varied with sexual swelling size, suggesting that calls signal female fertility. Possibly, Guinea baboon copulation calls represent an evolutionary remnant, no longer under selective pressure, and can be considered index signals of female fertility.

3
Insights for Estimating Animal Movement Step Selection Functions

Koshute, P.; Fagan, W. F.

2026-08-31 ecology 10.64898/2026.08.29.748012 medRxiv
Top 1%
0.1%
Show abstract

Ecologists remotely track movement steps of animals (e.g., via global positioning systems) and use step selection functions to study the effect of environmental factors upon their movement decisions. Constructing such functions requires pairing each observed step with some number of unobserved but feasible comparison steps. Larger numbers of comparison steps generally yield better estimates but also incur potentially challenging computational demands. Thus, it is important to determine an appropriate number of comparison steps. No established guidance exists for this decision. Here, we use simulated tracks to assess how many comparison steps are needed, fitting each set of steps to a conditional logistic regression model. We monitor errors in estimated effects for several classes of tracks, identifying the number of comparison steps for which mean relative absolute error in estimated effects is consistently low. By this criterion, 32 comparison steps per observed step are needed for our primary class of simulated tracks. Tracks in more homogeneous landscapes, tracks with shorter mean step lengths, or shorter tracks generally require more comparison steps (ranging from 64 to 128 per observed step) to achieve the same level of accuracy. Longer tracks generally require fewer comparison steps (16 per observed step). These results clearly demonstrate that the number of comparison steps influences how well step selection functions estimate covariate effects and provides initial direction in a research area that currently lacks quantitative guidance. Movement ecologists should take care when selecting the number of comparison steps paired with each observed step because those decisions matter.

4
Burden of fatigue in compensated chronic liver disease: findings from the multinational a:GAP Study

Choudhuri, G.; Akhundova-Unadkat, G.; Naidoo, N.; Morales-Castillo, M.; Guillaume, X.; Duijnhoven, R. G.; Safaei, A.; Swain, M. G.

2026-09-02 gastroenterology 10.64898/2026.08.28.26361618 medRxiv
Top 2%
0.0%
Show abstract

Background & Aims: Fatigue is a central symptom of chronic liver disease (CLD), substantially impacting health-related quality of life (HRQoL). This study aimed to further understand CLD symptomatology, including fatigue, and its impact on HRQoL from a patient perspective. Methods: Abbott Global Assessment of Patients unmet needs (aGAP) was a multinational, cross-sectional survey in adults with compensated CLD in China, India and Mexico, conducted between July and November 2024. Adult participants who self-reported that they had physician-diagnosed CLD and were experiencing fatigue completed a quantitative survey to assess symptom burden and included three HRQoL patient-reported outcome (PRO) questionnaires (Patient-Reported Outcomes Measurement Information System [PROMIS]-29+2, Work Productivity and Activity Impairment - Specific Health Problem version 2.0 [WPAI: SHP], Multidimensional Fatigue Inventory [MFI]). Results: Overall, 505 participants (China: 200; Mexico: 105; India: 200) completed the study. Participants reported that their CLD-related fatigue sometimes, often or always affected their self-esteem/confidence (45.1%) and ability to maintain or acquire new employment (38.6%). Most participants reported moderate (51.3%) or serious (26.9%) fatigue, with 33.5% experiencing fatigue every day or almost every day. Many participants felt their social life was negatively impacted by their fatigue (47.3%) and that there were related financial difficulties (53.9%). Use of validated PRO tools demonstrated severe fatigue (MFI: overall mean [SD] 13.9 [3.4] general fatigue and 13.4 [3.6] physical fatigue) as well as substantial levels of work and activity impairment (WPAI: SHP overall mean [SD] 53.0 [26.4]) and high levels of anxiety, pain interference, depression and sleep interference (PROMIS T-scores [≥]54). Conclusions: Fatigue has a substantial impact on HRQoL among adults with CLD across several countries, highlighting a global unmet need for targeted interventions to effectively identify and manage the condition.

5
The accuracy of urine-based mycobacterial antigens to detect childhood tuberculosis using an ultrasensitive immunoassay

Nkereuwem, E.; Misaghian, S.; Jaganath, D.; Calderon, R. I.; Luiz, J.; Paradkar, M.; Wambi, P.; Castro, R.; Nerurkar, R.; Wang, M.; Wohlstadter, J.; Franke, M. F.; Kampmann, B.; Kinikar, A.; Zar, H. J.; Segal, M.; Kato-Maeda, M.; Collins, J. M.; Swaney, D.; Cattamanchi, A.; Ernst, J. D.; Wobudeya, E.; Sigal, G.; The Combo Study,

2026-09-02 infectious diseases 10.64898/2026.08.28.26361530 medRxiv
Top 2%
0.0%
Show abstract

Background. Urine-based testing offers a promising non-sputum approach for diagnosing paediatric tuberculosis. However, the currently available lipoarabinomannan (LAM) assay shows limited sensitivity in children and is primarily indicated for those living with HIV. Co-detection of LAM with Mycobacterium tuberculosis (Mtb) proteins in urine could provide complementary pathogen-derived biomarkers that improve diagnostic performance. Methods. We developed an ultrasensitive multiplex electrochemiluminescence (ECL) immunoassay to measure Ag85B, CFP-10, ESAT-6, MPT32, and MPT64 in urine. We determined the analytical limits of detection and evaluated the diagnostic performance of individual proteins and LAM using urine samples from children with Confirmed, Unconfirmed, and Unlikely pulmonary tuberculosis enrolled across five high-burden countries (The Gambia, India, Peru, South Africa, and Uganda). Performance was assessed overall, by HIV and nutritional status, and across biomarker combinations. Findings. Urine samples from 630 children were analysed (median age was 4 years [IQR 2-8]; 44% female, 15% living with HIV, 19% underweight, 24% with Confirmed tuberculosis). The ECL assay achieved femtomolar limits of detection (1.5 to 4.0 fM). The sensitivity and specificity of individual Mtb proteins were 12-33% and 98-100%, respectively. Ag85B had the highest sensitivity (33%, 95% CI 26-41) for Confirmed tuberculosis and was similar to LAM. A four-antigen signature (Ag85B, MPT64, MPT32, LAM) was 50% sensitive (95% CI 42-58) and 94% specific (95% CI 90-96), and was significantly more sensitive than LAM alone, in particular among those without HIV. An additional sixteen (10%) of children with Unconfirmed TB had at least one Mtb protein or LAM detected. Interpretation. Multiple Mtb proteins are detectable in paediatric urine with high specificity, and multi-antigen signatures can augment sensitivity versus LAM alone. These findings demonstrate the potential of multi-antigen urine detection for childhood TB and define analytical targets for the development of future point-of-care diagnostics. Funding. National Institutes of Health.

6
Social Determinants of Health in HIV/HBV Coinfection Compared with HIV and HBV Monoinfection: A Framework for Dynamic Social Vulnerability

Yendewa, G.; Chengsupanimit, T.; Dehghani, A.; Ahmed, A.; Mohareb, A.; Freeman, M.; Cohen, C.; Ofotokun, I.; Dube, K.

2026-09-02 hiv aids 10.64898/2026.08.31.26361856 medRxiv
Top 2%
0.0%
Show abstract

Human immunodeficiency virus (HIV) and hepatitis B virus (HBV) coinfection is associated with accelerated liver disease, but whether coinfection is associated with newly documented social determinants of health (SDoH) is unclear. We conducted a retrospective cohort study using TriNetX across 110 U.S. healthcare organizations (2010-2026). We propensity score matched adults with HIV/HBV to adults with HIV or HBV monoinfection. We organized newly documented SDoH indicators using a dynamic individual-level framework with four clinically recognized domains of social disadvantage: material vulnerability, healthcare access and engagement, interpersonal adversity, and psychosocial vulnerability. Matched cohorts included 10,071 HIV/HBV-HIV pairs and 9,659 HIV/HBV-HBV pairs (mean age, 47 years; 79% male; 66% non-White; median follow-up, 3.3 years). Over 178,900 person-years, HIV/HBV was associated with higher risk of the primary SDoH composite compared with HIV (11.5% vs 9.7%; incidence rate, 2.50 vs 1.97 per 100 person-years; hazard ratio [HR], 1.25; 95% confidence interval [CI], 1.15-1.37) and HBV (11.0% vs 6.4%; incidence rate, 2.39 vs 1.67; HR, 1.50; 95% CI, 1.35-1.67). HIV/HBV was also associated with higher material vulnerability and healthcare access and engagement composites in both comparisons, including housing instability, food insecurity, financial insecurity, insurance instability, and care disengagement/nonadherence (HR range, 1.22-3.33 vs HIV; 1.31-1.94 vs HBV). In the HBV comparison, HIV/HBV was additionally associated with interpersonal adversity, primary support stressors, and violence or victimization (HR range, 1.36-2.16). Findings were robust across sensitivity analyses. HIV/HBV was associated with more newly documented SDoH than monoinfection, supporting dynamic SDoH assessment.

7
Evaluating Mean Platelet Volume in relation to Disease Severity in Paediatric Sickle Cell Anaemia: A Cross-Sectional Study in Kwara State, North-Central Nigeria

Oladimeji, F. D.; Adewoyin, A. D.; Oyeleke, K. O.

2026-09-02 hematology 10.64898/2026.08.28.26361349 medRxiv
Top 2%
0.0%
Show abstract

Background: Sickle cell anaemia (SCA) is characterised by chronic haemolysis, inflammation, platelet activation, and recurrent vaso-occlusive complications. Mean platelet volume (MPV) is a readily available platelet index, but evidence regarding its relationship with disease severity in paediatric SCA remains limited and inconsistent, particularly in African populations. Objective: To evaluate the relationship between MPV and disease severity among children with SCA in Kwara State, North-Central Nigeria. Methods: This hospital-based cross-sectional study included 51 clinically stable children with confirmed SCA consecutively recruited from the paediatric haematology clinic of Children Emergency Specialist Hospital, Ilorin. Complete blood count, including MPV, was performed using a Rayto RT-7600 automated haematology analyser. Disease severity was assessed using a composite clinical and laboratory scoring system based on a previously described method. Pearson's correlation, Spearman's rank correlation, simple linear regression, and the Kruskal-Wallis test were used as appropriate. Statistical significance was set at p < 0.05. Results: Of 51 participants, 14 (27.5%) had mild, 33 (64.7%) moderate, and 4 (7.8%) severe disease. Mean MPV was 9.34 +/- 0.76 fL (range, 8.0-11.2). Pearson's correlation showed a weak positive, non-significant linear relationship with severity score (r = 0.231, p = 0.103), whereas Spearman's analysis showed a weak positive monotonic association (rho = 0.286, p = 0.042). Regression explained 5.3% of severity-score variation (R2 = 0.053, p = 0.103). MPV did not differ significantly across severity categories (H = 2.163, p = 0.339). MPV correlated inversely with haemoglobin (r = -0.556, p < 0.001) and positively with platelet count (r = 0.307, p = 0.029). Conclusion: MPV showed a weak relationship with disease severity but inconsistent statistical evidence across analyses. The limited explained variance and absence of significant differences between severity categories do not support MPV as a standalone severity marker. Larger longitudinal studies are warranted. Keywords: Sickle cell anaemia; Mean platelet volume; Disease severity; Platelet indices; Paediatric haematology; Cross-sectional study; Nigeria.

8
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation

Chowdhury, A. R.; Chowdhury, B.

2026-09-02 health informatics 10.64898/2026.09.01.26361908 medRxiv
Top 2%
0.0%
Show abstract

Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.

9
ECG-based longitudinal risk prediction across diseases and organ systems

ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.

2026-09-02 health informatics 10.64898/2026.08.29.26361697 medRxiv
Top 2%
0.0%
Show abstract

Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.

10
Certified large language model-based diagnostic decision support in rheumatology: the ALLIANCE multicentre randomised controlled trial

Kremer, P.; Schlicker, N.; Hasnaj, R.; Bamberger, J.; Witte, T.; Haase, I.; Mayr, A.; Schmidt, C.; Osteras, N.; Baraliakos, X.; Kuhn, S.; Krusche, M.; Knitza, J.

2026-09-02 rheumatology 10.64898/2026.08.29.26361715 medRxiv
Top 2%
0.0%
Show abstract

Objectives To evaluate whether access to a certified large language model (LLM)-based clinical decision support system improves physician diagnostic performance in rheumatology compared with conventional diagnostic resources alone. Methods In this multicentre, open-label, randomised controlled trial, 82 physicians from seven hospitals in two countries were randomised 1:1 to conventional diagnostic resources plus Prof. Valmed or conventional resources alone. Participants assessed three rheumatology vignettes before and after assistance. The primary outcome was top-1 diagnostic accuracy. Secondary outcomes included top-3 accuracy, diagnostic reasoning, confidence, case-processing time and perceived support quality. Results Top-1 accuracy increased from 22.2% to 33.3% in the intervention group and from 23.3% to 35.0% in the control group, with no between-group difference in improvement (adjusted OR 0.99, 95% CI 0.45 to 2.19; p=0.979). Differences in top-3 accuracy, diagnostic reasoning and confidence were also not significant. Assisted case-processing time was substantially shorter with LLM support (94 vs 206 s; adjusted mean difference -112 s, 95% CI -141 to -83; p<0.001). Information timeliness and perceived diagnostic support quality were rated significantly higher in the intervention group. Exploratory analyses showed persistent overconfidence and substantial AI over-reliance. Conclusions Certified LLM-based diagnostic support did not improve diagnostic accuracy compared with conventional resources, but substantially reduced case-processing time and improved perceived support quality. These findings suggest potential workflow benefits while highlighting overconfidence and over-reliance as important safety considerations.

11
Optimizing Aqueous Humor Liquid Biopsy: Safety and Performance of a Short, Low-Dead-Space Ophthalmic Needle for Anterior Chamber Paracentesis

Singh, A. M.; Yeh, T.-C.; DeBoer, C.; Al-Moujahed, A.; Lin, J. B.; Smith, S. J.; Sanislo, S.; Janjua, K. A.; Lin, T.-C.; Almeida, D. R. P.; Mruthyunjaya, P.; Mahajan, V. B.

2026-09-02 ophthalmology 10.64898/2026.08.26.26361364 medRxiv
Top 2%
0.0%
Show abstract

Purpose: To evaluate the safety, procedural performance, sample recovery, and surgeon preference of an ophthalmic needle designed specifically for anterior chamber (AC) paracentesis. Methods: In this multicenter study, AC paracentesis was performed in clinic and operating-room settings using a 32-gauge x 4-mm needle with low dead space. The procedure was evaluated using a standardized physician survey. Prespecified outcomes included procedure-related adverse events (primary outcome), needle entry and handling, aspiration and sample recovery, comparative performance versus a 30-gauge needle, and physician preference for future use. Results: A total of 110 needle uses by eight surgeons were included. No ocular complications occurred, including lens or iris injury, hyphema, AC collapse, wound leak, hypotony, infection, or retinal complication, and no procedure required needle exchange or conversion to another device. Two technical events without ocular sequelae were noted, in which needle entry was partial thickness and did not reach the AC (1.8%; exact 95% CI, 0.2%-6.4%). Physicians rated needle entry, handling and sample recovery as good or excellent. Compared with a 30-gauge needle, the study needle was rated as at least comparable across all assessed domains. All surgeons rated it better or much better for intra-procedural safety and preferred it for future AC taps. Conclusions and Relevance: This short, 32-gauge low-dead-space ophthalmic needle demonstrated a favorable safety profile and was preferred over a 30-gauge needle by all surgeons. As aqueous humor liquid biopsy expands in clinical diagnostics and trials, an ophthalmic-specific needle design may help improve the consistency and safety of aqueous humor collection for molecular analysis and broader clinical use. Keywords: Anterior chamber paracentesis; Aqueous humor; Liquid biopsy; Low dead space; Ophthalmic needle

12
Evaluation of the Efficacy and Safety of Combination Therapy of Vamha and Myrha in the Management of PMOS: An Open-Label, Randomized, Multicentre, Comparative, Prospective Clinical Study

Patil, A.; Barathe, R.; Tate, D. M.; Kate, K.; Pande, S.; Gawande, N.; More, A.; Mahadik, S.; Berde, K.; Singhvi, R.

2026-09-02 sexual and reproductive health 10.64898/2026.08.20.26360875 medRxiv
Top 2%
0.0%
Show abstract

Introduction: Polyendocrine metabolic ovarian syndrome (PMOS), formerly known as polycystic ovary syndrome (PCOS), is a common endocrine disorder affecting women of reproductive age. Besides reproductive and metabolic disturbances, PMOS negatively impacts psychological well-being and quality of life. Despite available treatment options, there remains a need for safe and effective therapies that improve both clinical symptoms and fertility outcomes. Aim: To compare the efficacy of VAMHA and MYRHA tablet combination therapy with standard non-hormonal therapy in restoring regular menstruation. Secondary objectives included assessment of ovulation, menstrual symptoms, polycystic ovarian morphology, hormonal and metabolic parameters, anthropometric measures, and skin manifestations. Study Design: Open-label, randomized, multicentre, prospective comparative clinical study. Methods: Seventy-one women with PMOS were randomized to Group A (n=37) or Group B (n=34). Group A received VAMHA and MYRHA tablets (2 tablets each), while Group B received Metformin 500 mg plus Myoinositol 600 mg (1 tablet), twice daily for 180 days. Data were recorded in Case Report Forms. Statistical Analysis: Continuous variables were summarized using mean and standard deviation, while categorical variables were expressed as frequencies and percentages. Appropriate statistical tests, including Chi-square, were used. A p-value [&le;]0.05 was considered significant. Results: Significantly more participants in Group A achieved regular menstrual cycles than Group B (31 vs. 22; p<0.05). Ovulation occurred in 16 participants in Group A compared with 6 in Group B (p<0.05). Both groups showed significant improvement in menstrual irregularity and related symptoms. Significant reductions in Anti-Mullerian Hormone (AMH), fasting insulin, and body mass index (BMI) were observed in both groups (p<0.05). Resolution of polycystic ovarian morphology occurred in 13 participants (38.23%) in Group A and 10 (33.33%) in Group B. Both treatments were well tolerated with no major safety concerns. Conclusions: VAMHA and MYRHA combination therapy was superior to standard non-hormonal therapy in improving menstrual regularity and ovulation. It also produced favourable metabolic, hormonal, and ultrasonographic outcomes, suggesting its potential as a safe and effective option for comprehensive PMOS management and fertility enhancement.

13
GLP-1/GIP Uptake, Indication, and Access Pathways Among US Adults in the Understanding America Study

Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.

2026-09-02 endocrinology 10.64898/2026.08.28.26361368 medRxiv
Top 2%
0.0%
Show abstract

Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.

14
Can Dental AI Really Beat Dentists? DentalPair-Cert for Rigorous AI-Dentist Inference

Alve, S. R.; Rahman, S.; Meem, S. M. A. C.

2026-09-02 dentistry and oral medicine 10.64898/2026.09.01.26361874 medRxiv
Top 2%
0.0%
Show abstract

A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.

15
An interpretable, formally verified point-of-care ultrasound risk equation for difficult videolaryngoscopy: development and internal validation

Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.

2026-09-02 anesthesia 10.64898/2026.08.28.26361621 medRxiv
Top 2%
0.0%
Show abstract

Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.

16
Glaucoma and Diabetes Mellitus: A Comparative Evaluation of Comorbid Effect on Tear Quantity among Patients in Owerri, Imo State, Nigeria.

Chukwuoha, C. M.; Ovenseri-Ogbomo, G.; Azuamah, Y. C.; Odimegwu, N. E.; Obioma-Elemba, J. E.; Ugwoke, G.; Nkeremuzor, E. C.; Eronini, Y.; Ikoro, N. C.; Esenwah, E. C.

2026-09-02 ophthalmology 10.64898/2026.08.30.26361782 medRxiv
Top 2%
0.0%
Show abstract

Abstract Objective: Glaucoma is a chronic disorder that impairs ocular health and may exacerbate ocular surface disease leading to tear film instability, dry eye symptoms and decreased quality of life. This study compared changes in tear quantity among glaucoma subjects living with and without diabetes mellitus, attending an eye clinic in Nigeria. Methods: A comparative cross sectional research design was used. 157 subjects which comprised 74 glaucoma subjects living with diabetes mellitus and 83 glaucoma subjects living without diabetes mellitus participated in the study. Tear quantity assessment included the Schirmer I test and tear meniscus height (TMH) measurement. Descriptive statistics, independent samples t-test and Chi-square test were used to examine the data at 0.05 level of significance. Results: Glaucoma subjects living with diabetes mellitus showed substantially decreased tear production (11.4 +/- 6.8 mm) compared with glaucoma subjects living without diabetes mellitus (19.6 +/- 9.6 mm; p < 0.001). Tear meniscus height in glaucoma subjects living with diabetes mellitus (0.8 +/- 0.3 mm) was significantly greater than in subjects living without diabetes mellitus (0.7 +/- 0.3 mm; p = 0.034). Conclusion: Diabetes mellitus dramatically deteriorates the ocular surface function in glaucoma subjects by decreasing tear production, altering the tear meniscus height and increasing the severity of ocular surface symptoms. Routine glaucoma care, especially in patients with diabetes mellitus, should include a full ocular surface evaluation including Schirmer I test, TBUT, TMH, and OSDI assessment to allow early detection and management of ocular surface disease, better treatment adherence, and improved visual outcomes. Keywords: Glaucoma, Diabetes Mellitus, Tear production, Tear Meniscus Height, Ocular Surface Disease.

17
Non-inferior survival and enhanced longevity with initial low-dose versus full-dose enzalutamide: a single-centre real-world prostate cancer study

Gorobets, O.; Vinh-Hung, V.

2026-09-02 oncology 10.64898/2026.08.28.26361616 medRxiv
Top 2%
0.0%
Show abstract

Background: Prostate cancer enzalutamide treatment is approved at a standard dose of 160 mg daily. Concerns for real-world patients -- older and more fragile than those enrolled in clinical trials -- have prompted consideration of initiating treatment with lower doses, but the long-term efficacy of this approach remains unknown. We evaluate the long-term survival and longevity in patients treated with standard versus upfront low-dose enzalutamide. Methods: Retrospective analysis of 151 patients treated with enzalutamide (102 receiving 160 mg; 49 receiving [&le;]80 mg) between 2014--2021 at the Centre Hospitalier Universitaire de Martinique, with complete follow-up through end of life (98.7% completeness of follow-up). Primary outcomes were overall survival (OS), progression-free survival (PFS), and longevity (attained age). Results: Doses [&le;]80 mg were associated with longer median OS (36.3 vs. 20.7 months), improved restricted mean OS (difference of 0.7 years, p=0.05), and enhanced longevity (median 82.5 vs. 78.3 years, p=0.004). PSA response rate at 12 weeks was higher with lower-dose (71.4% vs. 48.8%, p=0.016). In multivariable models adjusted for prognostic factors, [&le;]40 mg compared with 160 mg was non-inferior regarding OS (HR=0.61, 95% CI 0.36--1.06), superior regarding PFS (HR=0.59, 95% CI 0.35--0.99), and superior regarding longevity (HR=0.48, 95% CI 0.28--0.84). Bone metastasis, poor performance status, PSA response, time to PSA nadir, and disease duration were independent predictors of outcomes. A post-hoc analysis revealed a strong association between dose and physician-prescribing profiles, ranging from "endorse-lowest-dose" to "never-deviate-from-full-dose". Conclusions: Lower doses of enzalutamide were non-inferior to full-dose. Dose-adapted strategies warrant further investigation.

18
Maternal cell-free RNA versus combined screening for first-trimester prediction of early-onset preeclampsia: a nested case-control study

Satorres-Perez, E.; Castillo-Marco, N.; Igual, M.; Cordero, T.; Munoz-Blat, I.; Monfort-Ortiz, R.; Marcos-Puig, B.; Simon, C.; Garrido-Gomez, T.; Perales-Marin, A.

2026-09-02 obstetrics and gynecology 10.64898/2026.08.28.26361628 medRxiv
Top 2%
0.0%
Show abstract

Background. In Europe, first-trimester combined screening with the Fetal Medicine Foundation (FMF) algorithm identifies women at increased risk of preeclampsia who may benefit from personalized aspirin prophylaxis. However, a substantial proportion of early-onset preeclampsia (EOPE) remains undetected at clinically acceptable specificity. Objective. To evaluate the first-trimester performance of MaiRa for early-onset preeclampsia (EOPE) risk stratification by benchmarking it against FMF screening in the same women, characterizing discordant patient-level classification profiles and exploring potential implementation strategies. Study Design. This secondary case-control analysis was nested within the prospective, multicentre PREMOM cohort [NCT04990141], which enrolled women with singleton pregnancies across 14 tertiary hospitals in Spain. First-trimester MaiRa and FMF risk estimates were evaluated in the same 126 pregnant women, comprising 99 uncomplicated controls and 27 EOPE cases, defined by disease onset before 34 weeks. Discrimination was compared using a stratified paired bootstrap analysis of the areas under the receiver-operating-characteristic curves. Performance was assessed at prespecified clinical thresholds, and detection rates were evaluated at fixed false-positive rates. Universal and contingent MaiRa implementation strategies were also evaluated. Results. MaiRa showed greater first-trimester discrimination for EOPE than FMF combined screening (AUC, 0.974 vs 0.900; P=.040) and consistently achieved higher detection rates across fixed false-positive rates. At false-positive rates of 5% and 10%, MaiRa detected 85.2% and 92.6% of EOPE cases, compared with 44.4% and 70.4% for FMF, respectively. Patient-level analysis demonstrated that MaiRa identified 12 of 27 EOPE cases (44.4%) classified as low risk by FMF; these pregnancies generally exhibited less abnormal conventional first-trimester profiles, including fewer maternal risk factors, lower mean arterial pressure and lower uterine artery pulsatility index, yet 8 of 12 (66.7%) subsequently developed severe EOPE. Exploratory implementation analyses showed that universal MaiRa screening achieved the highest EOPE detection, whereas a contingent strategy using FMF for triage and reflex MaiRa testing reduced molecular testing to 35.7% of pregnancies while maintaining 77.8% sensitivity and 97.0% specificity. Conclusion. MaiRa provided greater first-trimester discrimination for EOPE than conventional combined screening and detected additional pregnancies that later developed severe disease despite less abnormal conventional screening profiles. The findings suggest that maternal plasma cfRNA profiling captures biological alterations not fully reflected by combined first-trimester screening and support further prospective evaluation in an independent, unselected obstetric population. Key words: early-onset preeclampsia; first-trimester screening; cell-free RNA; liquid biopsy; Fetal Medicine Foundation algorithm; combined screening; risk stratification; aspirin prophylaxis.

19
Measuring Positive Stress Appraisal Among Nursing Students: Development and Psychometric Evaluation of the Nursing Student Positive Stress Scale (NSPSS)

Yan, H.; O'Brien, A. J.; Yoon, S. H.; Shaw, V.; vakavosaki, k.

2026-09-02 nursing 10.64898/2026.08.30.26361779 medRxiv
Top 2%
0.0%
Show abstract

Background: Stress research in nursing education has largely focused on distress, stressors, and negative outcomes, although challenging experiences may also support motivation, confidence, learning, and growth when appraised positively. Objective: To develop and evaluate the psychometric properties of the Nursing Student Positive Stress Scale (NSPSS). Design: A methodological instrument development and psychometric evaluation study. Methods: The NSPSS was developed using a deductive, theory-driven approach informed by the transactional theory of stress and coping and positive psychology perspectives. Content validity was assessed by an international nursing expert panel. Psychometric evaluation used national survey data from nursing students in New Zealand. Of 539 responses, 507 were analysed. Exploratory factor analysis (EFA) and confirmatory factor analysis (CFA) were conducted using separate subsamples. Internal consistency was assessed using Cronbach's alpha and McDonald's omega, and convergent validity through correlation with Perceived Stress Scale-10 scores. Results: Content validity was strong (I-CVI = .88-1.00; S-CVI/Ave = .975; S-CVI/UA = .800). EFA identified a dominant factor explaining 41.38% of variance (loadings = .528-.735). CFA supported a two-context Academic and Clinical Positive Stress model with correlated residuals between five parallel item pairs, chi-square(29) = 60.49, CFI = .970, TLI = .954, RMSEA = .063, SRMR = .065. Internal consistency was good (alpha = .839; omega = .843). NSPSS scores correlated negatively with PSS-10 scores (r = -.298, p < .001). Conclusion: The NSPSS demonstrated strong content validity, preliminary evidence of structural and convergent validity, and good internal consistency reliability for assessing positive stress appraisal among nursing students. Further validation in independent samples is warranted.

20
Cross-System Meta-Analysis of Machine Learning Predictors Identifies Value-Specific Risk Drivers and Interactions Underlying Acute Kidney Injury

Chan, H. Y.; Li, D.; Yu, A. S. L.; Kellum, J. A.; Fuhrman, D. Y.; Xu, Q.; Chrischilles, E. A.; Cowell, L. G.; Chandaka, S.; Anzalone, A. J.; Kean, J.; McTigue, K. M.; Mosa, A. S. M.; Taylor, B.; Syed, M.; Waitman, L. R.; Hu, Y.; Liu, M.

2026-09-02 nephrology 10.64898/2026.08.31.26361849 medRxiv
Top 2%
0.0%
Show abstract

Background: Current understanding of acute kidney injury (AKI) risk factors remains largely descriptive, offering limited precision into how specific biomarker values or physiologic thresholds influence susceptibility. We aimed to synthesize knowledge from machine learning models trained across multiple health systems to identify generalizable, value-specific risk drivers and biomarker interactions contributing to AKI risk. Methods: We analyzed electronic health records (EHRs) from 785,497 adult inpatients between 2010 and 2019 across nine U.S. academic medical centers within PCORnet. Interpretable gradient boosting machine models were independently developed at each health system to quantify predictor-outcome associations. Meta-regression was applied to integrate these site-level results, characterize nonlinear value-risk relationships, and identify bivariate interactions between predictors. Results: Meta-analysis revealed consistent, value-specific risk drivers across health systems. An increase in glucose from 100 mg/dL to 140 mg/dL was associated with a 1.46-fold higher risk of AKI. Chloride and anion gap also demonstrated elevated AKI risk with risk increases overlapping portions of their reference ranges, with anion gap showing a 1.14-fold increase across 4-12 mmol/L and chloride a 1.28-fold increase across 96-100 mEq/L. Electrolytes including potassium, calcium, and sodium showed quadratic associations with AKI risk. Bivariate meta-regression identified interactions between key predictors, highlighting pathways that jointly modulate AKI risk. Conclusion: This cross-system meta-analysis synthesizes machine learning-derived evidence into clinically interpretable knowledge, revealing how specific biomarker ranges and interactions modulate AKI risk. By moving beyond surface-level associations to quantitative, generalizable physiologic thresholds, these findings provide actionable insights to enhance risk stratification and personalized prevention in hospital care.