Schizophrenia
○ Springer Science and Business Media LLC
Preprints posted in the last 7 days, ranked by how well they match Schizophrenia's content profile, based on 21 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Mignondje, K. A.; Connolly, J. G.; Beermann, A.; Crabtree, E.; Vandekar, S.; Roeske, M. J.; Biernacki, K.; Coleman, M. J.; Shenton, M. E.; Brady, R. O.; Lewandowski, K. E.; Ward, H. B.
Show abstract
Background: Cognitive impairment is the leading cause of disability in schizophrenia with limited treatments. A major barrier to treatment development is the absence of reproducible, mechanistically grounded neural targets. Cross-sectional studies have identified dorsomedial prefrontal cortex (DMPFC)-somatomotor connectivity as a neural marker of cognitive performance on the Auditory Continuous performance task (ACPT), a measure of attention. To test the stability of this marker, we tested the relationship between DMPFC-somatomotor connectivity and ACPT performance in a longitudinal psychosis sample. Methods: Individuals with early psychosis (n=251) and matched controls (n=90) were enrolled and underwent resting-state neuroimaging and neurocognitive assessment. A subset completed longitudinal assessments over 2-4 years. We calculated DMPFC-somatomotor resting-state functional connectivity using a previously identified DMPFC region and a seed in the somatomotor cortex. We performed linear mixed effects models to predict ACPT performance based on connectivity, time, psychosis type, and their interaction. Results: In the psychosis sample, time (p=.0037) and affective psychosis diagnosis (p<.0001) predicted better ACPT performance. In a model predicting ACPT performance, we observed a significant interaction effect of DMPFC-somatomotor connectivity*psychosis subtype (p=.0079) such that DMPFC-somatomotor connectivity predicted ACPT performance only in individuals with non-affective psychosis (p=.0051). We then tested the specificity of this connectivity-cognitive performance relationship. In a model predicting DMPFC-somatomotor connectivity, only ACPT performance (p=.017), but not fluid cognition, was a significant predictor. Conclusions: DMPFC-somatomotor connectivity is longitudinally associated with cognitive performance in early psychosis. This relationship is strongest in nonaffective psychosis, suggesting a novel, reliable target for intervention for cognitive deficits in early psychosis.
Khan, Z.; McCarthy, C.; Dalton, K.; Jungo, K. T.; Doherty, A. S.; Reeve, E.; Moriarty, F.
Show abstract
Background: Adverse drug withdrawal events (ADWEs) are a key safety concern during deprescribing but remain poorly explored in pharmacovigilance systems. Objectives: To identify and compare ADWE signals across drug classes, different drugs within drug classes, and across patient characteristics, countries, and over time. Methods: A case/non-case disproportionality analysis was conducted in FDA-FAERS and EMA-EudraVigilance pharmacovigilance databases, with stratification by age (adults: 18-64, older adults: [≥]65), sex (male/female), reporting time (2004-2023 in 5-year intervals), and country (for EMA data). Disproportionality analysis (quantitative signal detection) was used to detect signals between ADWEs and drugs using the proportional reporting rate (PRR[≥]2), reporting odds ratio (ROR>1), and information component (IC>0) with case count [≥]5. Results: Overall, 158,501 reports (FDA-FAERS 145,514; EMA-EudraVigilance 12,987) included drug-event pairs related to ADWEs. In FDA-FAERS, clobetasone (IC=5.58; PRR=79.18; ROR=176.90) showed the strongest ADWE signals, followed by hydromorphone (4.85; 29.94; 37.37), hydrocodone, and paroxetine. In EMA-EudraVigilance, ethyl loflazepate (IC=6.01; PRR=119.80; ROR=197.53), clobetasone (5.39; 102.73; 155.10), veralipride, and levomethadone had the strongest signals. Most drugs maintained positive ADWE signals in analysis stratified into adults and older adults. However, among the top 10 drugs (based on highest IC values), buprenorphine/naloxone, desvenlafaxine, and baclofen in FDA-FAERS (ICs 4.95-6.05) showed stronger signals in older adults. A sex-based difference was observed, with paroxetine, venlafaxine, and buprenorphine/naloxone showing a stronger positive signal in females in both databases, whereas several opioids had stronger signals in males versus females across both databases. Conclusion: This study suggests ADWE signals for some medications differ by age and sex, potentially indicating different risks for withdrawal effects.
Wang, Y.; Zhang, E.; Guo, S.; Deng, A.; Xu, B.; Liao, J.; Wang, Y.; Dong, D.
Show abstract
Psychosis has long been conceptualized as a disorder of disrupted hierarchical integration across distributed brain systems, yet it remains unclear whether alterations in macroscale cortical hierarchy are already present before illness onset and are associated with subsequent transition to psychosis. Using connectome gradient mapping, we characterized baseline cortical hierarchical architecture along the unimodal-to-transmodal axis in 580 participants from the NAPLS-3 cohort, including converters (CHR-C, n = 56), non-converters (CHR-NC, n = 434), and healthy controls (HC, n = 90). Group differences were assessed at regional, network, and global levels. Group comparisons revealed that CHR-C individuals, relative to the other two groups, exhibited bidirectional alterations selectively along the sensorimotor-to-association gradient, with reduced values in the visual network alongside elevated values in the default mode network, indicating greater separation between sensory and transmodal systems along the gradient. At the global level, CHR-C showed increased explained variance, range, and variation of this gradient, collectively indicating hierarchical expansion. Notably, greater explained variance of this gradient was associated with a shorter time to conversion to psychosis, while increased gradient range and variation were associated with higher positive symptom severity across CHR individuals. These findings indicate that expansion of the sensorimotor-to-association connectome hierarchy is already present before psychosis onset in individuals who subsequently convert to psychosis. This altered hierarchical organization may reflect greater decoupling between sensory and transmodal systems and may characterize neurobiological changes associated with progression from a clinical high-risk state to psychotic illness.
Rohd, S. B.; Thorup, A. A.; Wilms, M.; Schiavon, M.; Streyma, D. H. B.; Laursen, A. F.; Bundgaard, A. F.; Sondergaard, A.; Krantz, M. F.; Veddum, L.; Hjorthoj, C.; Greve, A.; Mors, O.; Nordentoft, M.; Hemager, N.; Gregersen, M.
Show abstract
Objective: This study examined the prevalence of psychotic experiences (PE) and how early onset and persistence of PE contribute to risk and severity of mental disorders in adolescents at familial high-risk of schizophrenia (FHR-SZ) or bipolar disorder (FHR-BP) and adolescents from a population-based control group (PBC). Methods: This is the second follow-up of a nationwide cohort study including 522 children at FHR-SZ (N=202), FHR-BP (N=120), and PBC (N=200). Participants were assessed at ages 7, 11, and 15 using a semi-structured interview to evaluate PE and mental disorders. Results: At age 15, adolescents at FHR-SZ reported more PE than PBC over the past six months (current) and the past four years, while adolescents at FHR-BP only reported more current PE. PE reported at two or three timepoints (persistent PE) predicted any Axis I disorder in mid-adolescence, corresponding to three- (OR 2.9, 95% CI [1.5-5.7]) and 21-fold (OR 21.4, 95% CI [2.8-162.3]) increased risks, respectively. Persistent PE also predicted multimorbidity, with three- (OR 2.8, 95% CI [1.0-7.6]) and four-fold (OR 4.1, 95% CI [1.2-14.1]) increased risks, respectively. This was after adjustment for sex, early mental disorders, and familial risk. Conclusions: This study demonstrates a strong link between persistent PE and mid-adolescence mental disorders. Our findings emphasize PE as important risk markers for mental disorders during mid-adolescence and highlight the importance of monitoring children with PE before age 7 who develop persistent symptoms.
Quigley, H.; Gardiner, B.; McDaid, L.; O'Donnell, C.
Show abstract
Autism Spectrum Disorder (ASD) is a heterogeneous neurodevelopmental condition defined by differences in social communication and restricted, repetitive behaviours. As diagnostic criteria have broadened, ASD is now recognised across a wider range of individuals, raising key questions about its structure: does ASD have discrete sub-types, or is it better conceptualised as a continuous, possibly multidimensional, condition? We aim to explore whether a multidimensional continuum model more accurately captures the variability within ASD. We analysed a large SPARK phenotypic dataset of medical history and diagnostic surveys (background history, SCQ, RBS-R; n=36,710 individuals). We apply and compare two traditional statistical approaches, Factor Analysis and Gaussian Mixture Models, with a modern machine learning technique, the Variational Autoencoder (VAE). VAEs reconstructed unseen test data with ~4-fold better accuracy than Factor Analysis, and ~8-fold better accuracy than Gaussian Mixture Models. We identified four stable latent factors across 100 independently trained VAEs. These four dimensions provide an individual behavioural profile that can be visualized using radar-plots, offering a compact way to compare profiles at the person level. Through further analysis, we found evidence for 3 overlapping clusters or subtypes of ASD identified within the 4D latent space. This work aims to inform new ways of modelling ASD using a VAE that will be able to discern between a continuum or a clustered output and that go beyond binary diagnosis, instead reflecting the complex range of trait profiles, with implications for personalised diagnosis and intervention.
Page, S.; Easey, K.; Sedgewick, F.; Rai, D.; Stergiakouli, E.
Show abstract
A body of research suggests that autistic individuals are less likely to drink alcohol than neurotypicals. However, emerging studies support a link between autism and alcohol use. This complex relationship is also reflected in studies that have examined the genetic overlap between the two traits. However, it is unclear whether there is a direct causal relationship between them. To explore this, we applied a combination of polygenic score and Mendelian randomisation analyses using publicly available genome-wide summary statistics and phenotypic measures of autism and alcohol consumption from UK Biobank. LD score regression analyses did not provide evidence of a genetic correlation between genetic liability for autism and drinks consumed per week (rg=-0.08; CI95%=-0.19, 0.03). Further, findings from polygenic score analyses did not support an association between genetic liability for autism and overall monthly alcohol intake. Univariable Mendelian randomisation analyses showed little evidence for a total effect of autism, attention deficit hyperactivity disorder (ADHD) or depression on overall monthly alcohol consumption. Multivariable Mendelian randomisation analyses also showed little evidence of a direct effect of autism on drinks per week when controlling for ADHD and depression. It is plausible that genetic liability for autism does not directly increase the amount of alcohol consumed but instead operates via commonly co-occurring difficulties in the autistic community. However, our findings may be due to methodological shortcomings, including weak instruments biasing effects towards to the null. Consequently, results should be interpreted with caution and further research conducted to address these issues.
Ghuman, D.; Achar, T.; Gambhirrao, D.
Show abstract
Background Alcohol-associated injury is a leading cause of emergency department (ED) utilization in the United States and a clinically important driver of preventable morbidity across the adult lifespan. Prior surveillance research has characterized how the rate and severity of alcohol-associated injury vary by patient age, but whether the seasonal timing of injury risk is equally predictable across age groups (a question directly relevant to the timing of clinical screening intensification and public health intervention) has not been formally tested. Methods We conducted a retrospective surveillance analysis of 45,876 alcohol-associated ED visits among adults aged 18 years and older, identified from the National Electronic Injury Surveillance System (NEISS), 2019-2025 (weighted national estimate: 2,092,319 visits), using the structured Alcohol_Involved indicator introduced into NEISS case abstraction in 2019. Patients were stratified by sex and five age groups (18-24, 25-34, 35-49, 50-64, and [≥]65 years). Single-harmonic cosinor (Poisson) regression was used to estimate the seasonal peak day of injury risk (acrophase) for each stratum. To assess reliability, we performed leave-one-year-out jackknife resampling (seven iterations per group), case-resampling bootstrap confidence intervals (1,000 iterations), and likelihood-ratio tests of seasonal-phase interactions. Results Peak injury timing differed significantly across age groups (X^2 [8] = 2356.2, p < .0001). Adults aged 25-64 years showed a highly reproducible early-to-mid-July peak, with jackknife estimates shifting [≤]14 days when any single study year was excluded. Adults aged [≥]65 years showed significant seasonal variation annually (all p < .0001, amplitude comparable to younger groups) but a pooled peak estimate that shifted by up to 100 days across jackknife iterations. Sex-stratified analyses revealed that this instability was driven entirely by females aged [≥]65 years (jackknife range: 332 days, peak consistently in late October through early January) rather than males aged [≥]65 (jackknife range: 31 days, peak consistently in early August). Hospital admission rates increased monotonically with age from 9.0% (18-24 years) to 31.8% ([≥]65 years). Conclusions Alcohol-associated injury follows a reproducible, calendar-stable summer seasonal pattern in adults aged 25-64 years. Among adults [≥]65 years, the previously reported temporal instability is concentrated in the female subgroup, whose seasonal injury risk does not converge on a fixed calendar window. These findings suggest that fixed-calendar prevention and screening strategies are well suited to working-age adults and older men, but older women may require a year-round, individually tailored approach. Keywords: Alcohol-related injury; Emergency department; Seasonality; Age factors; Sex differences; Injury surveillance; Cosinor analysis; Older adults
Monteseirin, K.; Mendez-Couz, M.; Rivas-Fernandez, M. A.; Conejo, N. M.
Show abstract
Children and adolescents with hearing loss frequently encounter reduced auditory access and delayed language development, factors that may influence the maturation of executive functions. This study examined developmental differences in planning, a core executive function, in 98 children and adolescents with hearing loss or normal hearing aged 7 to18 years using the Tower of London task. Compared to normal hearing peers, participants with hearing loss made more unnecessary moves and rule violations and initiated problem-solving more rapidly, suggesting reduced preplanning efficiency and increased impulsivity. These group differences were most pronounced in adolescents, who showed faster initiation and greater movement inefficiency than age-matched normal hearing participants. Within the hearing loss group, adolescents displayed higher accuracy and longer initiation times than children, reflecting developmental improvements despite persistent gaps relative to hearing peers. Language development age did not alter the main effects. Findings indicate that reduced early auditory and language access may contribute to differences in planning development, highlighting the need for targeted executive functions support in educational and clinical settings for youth with hearing loss.
Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.
Show abstract
Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.
Bastien, J.; Garcia, K.; Wallace, A. L.; Sullivan, R. M.; Hoh, E.; Wade, N. E.
Show abstract
Background: As cannabis policy changes in the United States, secondhand cannabis smoke (SCS) is increasingly common, including within families. However, prevalence of exposure and clinical correlates over time in adolescents are not fully understood. Objectives: (1) To estimate the prevalence of SCS and personal cannabis use in US-based teens exposed to SCS, and (2) examine the cognitive trajectories of adolescents exposed to SCS compared to non-exposed peers. Methods: Data from the Adolescent Brain Cognitive Development (ABCD) Study was used. Participants (n=11,316 of full cohort with follow-up data; n=776 with self-reported family SCS exposure) attended yearly visits from ages 11-17, completing substance use interviews, toxicological testing, and the NIH Toolbox Cognitive battery. Youth with SCS but no personal cannabis use (n=419; 47% female) were matched on prenatal substance exposure, family substance use history, and sociodemographics to non-SCS exposed and non-cannabis-using youth with a 1:2 ratio (Controls n=838). Linear mixed-effects models assessed cognitive performance by SCS*age interactions, accounting for random effects of subject and family. Covariates included sex and alcohol, nicotine, and other substance use. Secondary models analyzed performance by cumulative waves of reported SCS exposure interacting with age. Results: Of the full cohort, 6.9% (n=776) reported exposure to SCS. Of these individuals, 46% endorsed lifetime personal cannabis use by age 17, relative to 20% of non-SCS exposed youth (OR=3.83[95%CI:3.29,4.44]). Within matched participants, SCS*age demonstrated a significant interaction on attention and inhibitory control ({beta}=-0.32, p=.028), with SCS demonstrating reduced improvement over time. More waves of exposure were also associated with worse performance over time ({beta}=-0.39, p=.057). Discussion: Almost half of those who had been exposed to SCS endorsed personal cannabis use. Cognitive findings were domain specific, similar to findings in secondhand tobacco: SCS exposed youth showed restricted improvement in attention and inhibitory control by age 17. Public health and policymakers should make efforts to curb youth SCS exposure, given the potential for risk which has not been fully explored to date.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Sörnyei, D.; Kovacs, F. M.; Benedek, T.; Ori, D.; Farkas, K.
Show abstract
The Autism Spectrum Quotient (AQ-50) is widely used to assess autistic traits, yet its Hungarian version has not been psychometrically evaluated. We assessed the reliability, factor structure, temporal stability, convergent validity, and clinical utility of the Hungarian AQ-50 and a revised translation (AQ-50-HU-R) in two samples (N1 = 1967; N2 = 423), including autistic and non-autistic participants. The AQ-50-HU-R showed high internal consistency and test-retest reliability. A bifactor model provided the best fit ({chi}2[1125] = 1650.433, p < 0.001; CFI = 0.991; TLI = 0.990; RMSEA = 0.033 [90% CI = 0.030-0.037]; SRMR = 0.083), with 71% of common variance attributable to a general autistic traits factor. The total score distinguished clinically verified autistic participants from participants reporting no ASD diagnosis (AUC = 0.906), with a cutoff of 25. Associations with ADOS scores were weak or nonsignificant. The AQ-50-HU-R is best interpreted as a reliable total-score screening measure, supporting referral for comprehensive autism assessment.
Pham, T. M.; Smith, J. T.; Mortimer, T. D.; Grad, Y.; Earl, A. M.; Lewis, I. A.; PRIME Consortium,
Show abstract
Background Using a population-based cohort from the Calgary Health Zone (CHZ), Canada, we integrated longitudinal antimicrobial susceptibility and prescribing data with the whole genome sequences of five major pathogens. We aimed to assess how antimicrobial resistance (AMR) responds to prescribing changes and determine which bacterial strains shape these dynamics. Methods We analysed antibiotic prescribing rates, clinical and genomic data from 7,271 Staphylococcus aureus, 1,609 Enterococcus faecalis, 801 Enterococcus faecium, 11,363 Escherichia coli, and 2,319 Klebsiella pneumoniae isolates, associated with bacteraemia episodes in the CHZ between 2006-2022. Genomic clusters (referred to as strains) were identified using StrainGST and assigned to known sequence types (STs) or clonal complexes (CCs). Strain-level incidence, stratified by community-onset (isolates collected [≤]48h after admission) and hospital-onset (>48h after admission), AMR phenotypes, and prescribing rates were modelled using negative-binomial and binomial regression. Temporal trends were quantified using average annual percentage change (AAPC). Findings Between 2010-2022, fluoroquinolone prescribing declined in both community (AAPC=-6.8% [95% CI -8.1, -5.4]; p<0.0001) and hospital settings (AAPC=-5.1% [-6.5, -3.7]; p<0.0001). This was accompanied by a significant reduction in fluoroquinolone resistance among Gram-positive species. Specifically, S aureus bacteraemia resistant to clinically important antibiotics, cloxacillin, ciprofloxacin, erythromycin, and clindamycin, declined from 2006 to 2022, mostly in hospital-onset cases (AAPC=-16.0%, [-19.3%, -12.7%], p<0.0001). In E coli, ceftriaxone and ciprofloxacin resistance were clustered in ST131 and the emerging ST1193; the latter increased steadily, particularly in community-onset cases (AAPC=17.7%, [0.0%, 30.0%], p<0.0001). CTX-M-27-producing E coli ST131 strains increased (AAPC=23.8%, [17.4%, 30.5%], p<0.0001) between 20082022, while CTX-M-14-producing E coli ST131 declined (AAPC=-15.9%, [-21.3%, -10.2%], p<0.0001) between 2013-2022. These trends were paralleled by an increase in community cephalosporin prescribing (AAPC=7.3%, [4.2%, 10.5%], p<0.0001) between 2010-2022. For K pneumoniae, hypervirulent ST23 was most common (N=88) with an increasing trend in incidence (AAPC=3.0%, [-2.8%, 9.2%]) between 2006-2019. Conclusions The contrasting resistance trends between Gram-positive and Gram-negative species underscore the complexity of AMR control efforts. Effective strategies will require stewardship efforts targeting multiple drug classes, genomic surveillance for emerging resistant strains, and interventions extending beyond hospital settings.
Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.
Show abstract
Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.
Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.
Show abstract
Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [≥]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.
Kronlage, C.; Ripart, M.; Piper, R. J.; Tisdall, M. M.; Carmichael, D. W.; Baldeweg, T.; Duncan, J. S.; O'Muircheartaigh, J.; Eriksson, M. H.; Casella, C.; Bridgen, P.; Bauer, T.; Bouschery, S. R.; Lange, A.; Pracht, E. D.; Stocker, T.; Surges, R.; Ruber, T.; Klodowski, K.; Rodgers, C. T.; Cope, T. E.; Wagstyl, K.; Adler, S.
Show abstract
Background: Hippocampal sclerosis (HS) is a common cause of drug-resistant focal epilepsy (DRFE) and amenable to neurosurgical treatment. Detection relies on MRI but can be challenging. 7 Tesla (T) ultra-high field MRI and automated MRI post-processing tools have independently been shown to improve radiological diagnosis of HS. However, combining these approaches remains underexplored. This study evaluated whether AID-HS, a tool for HS detection developed using 3T MRI, generalises to 7T MRI data. Methods: We collated a dataset of paired 3T and 7T T1-weighted MRI from four epilepsy centres, including 23 patients with HS, 39 healthy controls, and 23 individuals with focal cortical dysplasia as disease controls. Histopathology served as the gold standard for defining HS where available (n=7), otherwise radiological findings (n=16). AID-HS was applied to images acquired at both field strengths, and sensitivity and specificity for detection and lateralisation of HS were compared. Additionally, agreement of hippocampal features across 3T and 7T was evaluated. Results: We found no evidence of a difference in performance of AID-HS between 3T and 7T. Sensitivity for detection of unilateral HS was 63% (12/19) at 3T and 68% (13/19) at 7T (McNemar's exact test p=1.0). Specificity in controls was 97% (60/62) at 3T and 100% (62/62) at 7T (p=0.5). Bilateral HS was correctly flagged in 3 of 4 cases using feature-based criteria, with high specificity in controls. Quantitative hippocampal features showed moderate to good agreement across field strengths (ICC 0.70 to 0.98), with small differences observed for volume and thickness estimates. Conclusion: AID-HS provides robust detection and lateralisation of HS across multiple 7T MRI centres, highlighting its potential to enhance lesion detection. Future work is needed to investigate whether models trained on 7T data can leverage the improved image quality for further gains in HS detection performance.
Gorenshtein, A.; Omar, M.; Jia, E. L.; Adiniaev, Y.; Daniel, O.; Kruskal, J.; Ahmed, M.; Brook, O. R.; Klang, E.; Barash, Y.
Show abstract
Objective: Published P300-speller fusion schemes fix prior trust regardless of trial reliability; we tested whether a reliability estimate improves on it. Methods: We reanalyzed 3,373 archived P300-speller selections from 47 people with ALS (BigP3BCI). A fair, matched-search-space comparison, tuning both a fixed weight and an adaptive policy out-of-fold, was evaluated across 22 evaluable language-model priors up to 46.7B parameters. Two representative priors, GPT-2 and a classical 5-gram, additionally received detailed naive and mechanistic analyses. Results: No prior's 95% CI favored adaptive fusion under the fair comparison, despite unexploited oracle headroom at every scale. Under GPT-2, the naive comparison was significantly worse for adaptive fusion; both anchors converged to a degenerate or near-degenerate fair-comparison solution. For the representative anchors, three further controllers failed to convert that headroom into benefit; the fixed-fused posterior's output probability outperformed the best controller for flagging errors (2.8- to 3.8-fold enrichment). Conclusion: A tuned fixed weight is a difficult-to-beat default across the tested scale range; reliability estimation gave no deployable adaptive advantage. Significance: Adaptive weighting should be validated against a fairly tuned baseline across model families and scales; in this dataset, the fused output's confidence identified high-risk selections better than the tested purpose-built ranker.
Chen, P.-H.; Duncan, N. W.; Lee, H.-c.; Liu, Y.-J.; Hsu, T.-Y.
Show abstract
Background: Bipolar disorder is associated with persistent social, cognitive, and functional impairment during euthymia, yet the neural mechanisms underlying these deficits remain unclear. Alterations to self-referential processing are a candidate mechanism, but existing electrophysiological studies rely on emotionally valenced paradigms that potentially confound self-processing with emotional biases. Methods: We analysed electroencephalography from 28 patients with bipolar disorder (type I or II) and 28 age- and sex-matched healthy controls during an emotionally neutral colour judgment task with self-related (preference) and non-self-related (similarity) conditions. Late positive potentials, temporal generalisation decoding, and frequency band decoding (theta, alpha, beta) were used to characterise the temporal dynamics and oscillatory correlates of self versus non-self processing. Results: Controls showed higher overall event-related potential amplitudes and greater self versus non-self differentiation than patients (condition by group interaction, 337 to 946 ms). Broadband temporal generalisation decoding revealed extensive cross-temporal generalisation of the self versus non-self representation in controls, spanning most of the trial, but no significant generalisation in patients. Frequency analyses showed that alpha and beta carried self versus non-self information in both groups, with broader extent in controls, and that anterior theta carried this information in patients but not controls. Exploratory correlations linked decoding measures to rumination and anxiety but not to manic symptoms. Conclusions: The neural representation distinguishing self-referential from externally guided processing was both smaller in amplitude and less temporally sustained in bipolar disorder. Reduced persistence is not detectable by conventional amplitude analyses, and may bear on the self-related and social cognitive difficulties reported in this population.
Chesley, J.; Biernacki, K.; Vanleuven, J.; Doran, J. P.; Yazgan, I.; Yildiz, G.; Gonzalez, D. A.; Wagner, S. Y.; LeBaron, K.; Marrero, E.; Osama, T.; Vandekar, S.; Ward, H. B.
Show abstract
Background: Substance use is common among individuals with depression. Transcranial magnetic stimulation (TMS) is an effective treatment for depression, but current clinical guidelines have discouraged TMS treatment for individuals with depression and co-occurring substance use given concerns for limited efficacy. However, limited data exists on whether substance use affects response to TMS. Methods: Using electronic health record data from patients who received a standard course of TMS for major depressive disorder at an academic medical center, we investigated associations between substance use frequency and response to TMS, defined as change in Patient Health Questionnaire-9 (PHQ-9) scores. Substance use frequency was extracted for alcohol, cannabis, nicotine, stimulants, benzodiazepines, opioids, inhalants, psychedelics, and other drugs. We performed ANCOVA and multiple regression analyses to predict change in PHQ-9 score based on substance use frequency, controlling for pre-TMS PHQ-9 score, age, sex, and number of TMS sessions received. Results: We extracted data from 219 TMS courses. Alcohol was the substance used most commonly (34.2%), followed by prescription benzodiazepines (28.3%), and prescription stimulants (21.0%). Across all substance categories, substance use was not associated with change in PHQ-9 score (all p > 0.05, Cohens d=0.00 to 0.30). In multiple regression models to compare individual levels of substance use frequency (e.g., daily use vs. no use), level of substance use was not associated with change in PHQ-9 score (all p > 0.05). The range of plausible effects of substance use frequency on PHQ-9 change was generally below the minimal clinically important difference for PHQ-9, suggesting substance use was unlikely to have a meaningful clinical effect on antidepressant response to TMS. Conclusions: Low to moderate substance use does not have a clinically significant effect on antidepressant response to TMS. Low-level substance use should not exclude individuals with depression from receiving TMS.
Chen, Y.; Puckett, H.; Clarot, G.; Hawkins, B.; Sharp, K.; Todd, D. A.; Lopez, A.; Bertollo, J. R.; Behar, H. E.; Zeithamova, D.; Xie, H.; Verbalis, A.; VanMeter, A. S.; Gaillard, W. D.; Kenworthy, L.; Vaidya, C. J.
Show abstract
Generalization is a key cognitive process that allows humans to flexibly apply prior knowledge to guide new behaviors. Difficulties with generalization and flexibility are observed across neurodevelopmental disorders, especially autism, limiting adaptive function and quality of life. Cognitive-behavioral treatment benefits some but not all autistic individuals. As treatment requires application of learned skills to everyday life, variability in generalization ability may limit intervention success in autism. While cognitive substrates of learning and generalization are well established, their potential for explaining clinical outcomes is not known. Here, we combined a category learning task with computational modelling to distinguish two learning strategies underlying generalization -- prototype abstraction vs. exemplar memorization -- and tested whether individual differences in these learning strategies predicted real-world intervention outcomes in autistic youth. Fifty-four participants completed the category learning task at two pre-intervention timepoints, and then completed Unstuck and On Target:14-22 intervention targeting flexible problem solving, goal setting, and planning. We found that participants who consistently relied on prototype abstraction (N=26) were subsequently more likely to benefit from the intervention, showing improvement in parent- and self-reported flexibility. These findings identify prototype abstraction as a clinically relevant cognitive capacity that may help explain individual differences in intervention response and support the tailoring of interventions. More broadly, they demonstrate the value of linking basic cognitive mechanisms to clinical outcomes and may inform strategies to enhance the effectiveness of cognitive-behavioral interventions for youth with developmental disabilities.