Back

Poor prognosis of stage I lung adenocarcinoma patients determined by elevated expression over pre/minimally invasive status of COL11A1 and THBS2 in the focal adhesion pathway

Shang, J.; Zhao, Y.; Jiang, H.; Yang, J.; Zhang, N.; Ren, L.; Chen, Q.; Yu, Y.; Shi, L.; Chen, H.; Zheng, Y.

2021-12-17 oncology
10.1101/2021.12.16.21267913 medRxiv
Show abstract

Around 20% of stage I lung adenocarcinoma (LUAD) patients die within five years after surgery, and efforts for developing gene-expression based models for risk-tailored post-surgery treatment are largely unsatisfactory due to overfitting-related lack of validation and extrapolation. Because patients with adenocarcinomas in situ (AIS) and minimally invasive (MIA) LUAD are completely curable by surgical resection, we hypothesize that poor-prognosis stage I patients may exhibit key molecular characteristics deviating from AIS/MIA. We first found focal adhesion (FA) as the only pathway significantly perturbed at both genomic and transcriptomic levels by comparing 98 AIS/MIA and 99 invasive LUAD patients. Then, we identified two FA pathway genes (COL11A1 and THBS2) strongly upregulated from AIS/MIA to stage I while expressed steadily from normal to AIS/MIA. Furthermore, unsupervised clustering separated stage I patients into two molecularly and prognostically distinct subtypes (S1 and S2) based solely on the expression levels of COL11A1 and THBS2 (FA2). Subtype S1 looked like AIS/MIA, whereas S2 exhibited more somatic alterations, elevated expression of COL11A1 and THBS2, and more activated cancer-associated fibroblast (CAF). The prognostic performance of the knowledge-driven and overfitting-resistant FA2 model was validated with 12 external data sets and may help reliably identify high-risk stage I patients for more intensive post-surgery treatment.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.