World Journal of Oncology, ISSN 1920-4531 print, 1920-454X online, Open Access
Article copyright, the authors; Journal compilation copyright, World J Oncol and Elmer Press Inc
Journal website https://wjon.elmerpub.com

Original Article

Volume 17, Number 5, October 2026, pages 683-704


A Mitochondrial-Related Gene Signature for Diagnosis and Immune Microenvironment Modulation in Lung Cancer and Venous Thromboembolism

Figures

↓  Figure 1. Technology roadmap. LCDEGs: lung cancer-associated differentially expressed genes; VTEDEGs: venous thromboembolism–associated differentially expressed genes; CGs: crosstalk genes; WGCNA: weighted gene co-expression network analysis; GSVA: gene set variation analysis; GO: Gene Ontology; KEGG: Kyoto Encyclopedia of Genes and Genomes; LASSO: least absolute shrinkage and selection operator; TF: transcription factors; ssGSEA: single-sample gene set enrichment analysis.
Figure 1.
↓  Figure 2. Differential gene expression analysis in LC. (a, b) Volcano plots of DEGs in GSE30219 (a) and GSE19188 (b), comparing LC samples with controls. (c) Venn diagram showing the overlap of DEGs between GSE30219 and GSE19188. (d, e) Heatmaps of the top 10 upregulated and top 10 downregulated LCDEGs in GSE30219 (d) and GSE19188 (e). In volcano plots, orange indicates LC samples, and light blue indicates controls. In heatmaps, red indicates higher expression and blue indicates lower expression. DEGs: differentially expressed genes; LC: lung cancer.
Figure 2.
↓  Figure 3. WGCNA of the GSE30219 dataset. (a) Determination of the optimal soft-thresholding parameter to establish a scale-free network architecture. The left panel illustrates the scale-free topology fit index across varying powers, while the right panel depicts mean connectivity. (b) Hierarchical clustering dendrogram of genes derived from the top 50% most variable genes. (c) Module assignment based on hierarchical clustering. The upper section illustrates the gene dendrogram, while the lower section depicts the associated gene modules using distinct color codes. (d) Heatmap of correlations between gene modules and sample groups (LC vs. controls). Correlation strength is indicated by |r| values, where 0.3–0.5 represents weak associations and 0.5–0.8 represents moderate associations. Color indicates direction of association: pink indicates positive correlations and gray-green represents negative correlations. WGCNA: weighted gene co-expression network analysis; LC: lung cancer.
Figure 3.
↓  Figure 4. Gene Ontology enrichment analysis of crosstalk genes. (a) Bubble plot showing GO enrichment results for CGs across BP, CC, and MF. GO terms are shown on the x-axis. Bubble size indicates the number of genes, and color indicates statistical significance, with red indicating lower P-values and blue indicating higher P-values. (b–d) Network plots of enriched GO terms for CGs in BP (b), CC (c), and MF (d). Nodes represent GO terms (light red) and genes (light blue), and edges indicate their associations. Established enrichment thresholds were P < 0.05 and false discovery rate (FDR) < 0.25. CGs: crosstalk genes; GO: Gene Ontology; BP: biological process; CC: cellular component; MF: molecular function.
Figure 4.
↓  Figure 5. GSVA profiling of the GSE30219 LC dataset. (a) Heatmap of GSVA enrichment scores comparing LC samples with controls. Red corresponds to higher pathway activity; blue corresponds to lower activity. (b) Boxplots showing differences in pathway enrichment scores between the LC and control groups. Orange represents LC samples, and light blue represents controls. Pathway selection was conducted according to P < 0.05. ***P < 0.001, reflecting significant differences between groups.
Figure 5.
↓  Figure 6. Identification of pivotal diagnostic genes. (a) LASSO cross-validation curve showing model performance in the LC dataset (GSE30219). (b) Coefficient paths illustrating variable selection in the LC dataset (GSE30219). (c) Forest plot showing effect estimates and confidence intervals of diagnostic genes associated with LC. (d) Diagnostic accuracy assessment of the CG-derived LASSO model applied to the VTE dataset (GSE48000). (e) Trajectory plots depicting the shrinkage paths of coefficients for predictors in the VTE cohort (GSE48000) under LASSO regularization. (f) Forest plot showing effect estimates and confidence intervals of diagnostic genes associated with VTE. (g) Venn diagram showing the overlap between LC- and VTE-associated diagnostic genes identified by LASSO. LC: lung cancer; VTE: venous thromboembolism; LASSO: least absolute shrinkage and selection operator; CGs: crosstalk genes.
Figure 6.
↓  Figure 7. Development and validation of diagnostic logistic regression models using signature genes. (a) Nomogram for predicting LC based on the multivariable logistic regression model in GSE30219. (b) Calibration plot showing agreement between predicted and recorded findings in GSE30219. (c) DCA plot evaluating the practical application of the model in GSE30219 in clinical settings. (d, e) ROC curves showing model performance in the primary dataset GSE30219 (d) and the validation dataset GSE19188 (e). (f) Nomogram for predicting VTE utilizing a multivariable logistic regression framework in GSE48000. (g) Calibration plot showing consistency between anticipated and actual results in GSE48000. (h) DCA plot evaluating the practical application of the model in GSE48000. (i, j) ROC curves showing model performance in GSE48000 (i) and the validation dataset GSE19151 (j). In DCA plots, the y-axis quantifies net benefit, and the x-axis indicates threshold probability. In ROC curves, AUC values closer to 1.0 indicate superior diagnostic efficacy, values of 0.5–0.7 indicate suboptimal precision, and 0.7–0.9 indicate moderate precision. DCA: decision curve analysis; ROC: receiver operating characteristic; AUC: area under the curve; TPR: true positive rate; FPR: false positive rate.
Figure 7.
↓  Figure 8. Validation of differential expression and diagnostic performance of candidate genes. (a) Expression comparison of key genes between LC and control samples in GSE30219. (b, c) Further analyses in GSE30219: (b) expression levels of ACAA1 and HSD17B10; (c) ROC curves for MTIF2, THOP1, and PDE2A. (d) Expression comparison of key genes between LC and control samples in GSE19188. (e, f) Further analyses in GSE19188: (e) expression levels of ACAA1 and HSD17B10; (f) ROC curves for MTIF2, THOP1, and PDE2A. (g) Expression comparison of critical genes distinguishing VTE specimens from control cohorts in GSE48000. (h, i) Further analyses in GSE48000: (h) expression levels of ACAA1 and HSD17B10; (i) ROC curves for MTIF2, THOP1, and PDE2A. (j) Expression comparison of key genes between VTE and control samples in GSE19151. (k, l) Further analyses in GSE19151: (k) expression levels of ACAA1 and HSD17B10; (l) ROC curves for MTIF2, THOP1, and PDE2A. Statistical significance is denoted as follows: ns, not significant (P ≥ 0.05); **P < 0.01; ***P < 0.001. Diagnostic performance is interpreted based on AUC values: 0.5–0.7 reflects limited discriminatory power, 0.7–0.9 denote moderate performance, and > 0.9 indicates excellent accuracy. ROC: receiver operating characteristic; LC: lung cancer; AUC: area under the curve; VTE: venous thromboembolism. Color coding: blue, control; orange, LC; pink, VTE.
Figure 8.
↓  Figure 9. Investigation of regulatory networks governing key genes. (a) PPI network constructed using STRING, showing interactions among the five genes. (b) GeneMANIA network showing functional associations between the five genes and related proteins. Edge colors represent distinct interaction types. (c) mRNA–TF regulatory network showing interactions between genes and TFs. (d) mRNA–miRNA regulatory network showing interactions between genes and miRNAs. In panels c and d, orange nodes represent mRNAs, yellow nodes represent TFs, and purple nodes represent miRNAs. TF: transcription factor; PPI: protein–protein interaction.
Figure 9.
↓  Figure 10. Immune infiltration profiling using ssGSEA in the GSE19188 dataset. (a) Boxplots comparing the abundance of immune cell subsets between control and LC samples. (b) Heatmap showing correlations among immune cell populations in GSE19188. (c) Bubble plot illustrating the correlations between gene expression levels and immune cell infiltration. Statistical significance is indicated as follows: ns, not significant (P ≥ 0.05); *P < 0.05; **P < 0.01; ***P < 0.001. Correlation strength is based on |r|, where < 0.3 reflects negligible, 0.3–0.5 weak, 0.5–0.8 moderate, and > 0.8 strong correlations. Color coding: light blue, control; orange, LC. In panels b and c, red indicates a positive correlation and blue indicates a negative correlation, with color intensity reflecting correlation strength. ssGSEA: single-sample gene set enrichment analysis; LC: lung cancer.
Figure 10.

Tables

↓  Table 1. Results of GO Enrichment Analysis for CGs
 
ONTOLOGYIDDescriptionGeneRatioBgRatioP valueP.adjustq value
GO: Gene Ontology; BP: biological process; CC: cellular component; MF: molecular function; KEGG: Kyoto Encyclopedia of Genes and Genomes; CGs: crosstalk genes.
BPGO:0006753Nucleoside phosphate metabolic process6/37495/18,8000.000375510.046861290.03586955
BPGO:1901293Nucleoside phosphate biosynthetic process5/37257/18,8000.000139970.040871290.03128456
BPGO:0006163Purine nucleotide metabolic process5/37394/18,8000.000989120.051376130.0393254
BPGO:0072521Purine-containing compound metabolic process5/37416/18,8000.001259650.051376130.0393254
BPGO:0009117Nucleotide metabolic process5/37487/18,8000.002512830.063996640.04898565
CCGO:0005759Mitochondrial matrix15/39473/19,5946.5367 × 10−158.1055 × 10−136.5367 × 10−13
CCGO:0098798Mitochondrial protein-containing complex10/39281/19,5941.3819 × 10−108.5675 × 10−96.9093 × 10−9
CCGO:0005743Mitochondrial inner membrane9/39491/19,5943.9314 × 10−71.625 × 10−51.3105 × 10−5
CCGO:0000313Organellar ribosome4/3981/19,5941.9965 × 10−50.000454440.00036649
CCGO:0005761Mitochondrial ribosome4/3981/19,5941.9965 × 10−50.000454440.00036649
MFGO:0051087Chaperone binding3/39106/18,4100.001458350.068345940.04778702
MFGO:0019001Guanyl nucleotide binding5/39401/18,4100.001496630.068345940.04778702
MFGO:0032561Guanyl ribonucleotide binding5/39401/18,4100.001496630.068345940.04778702

 

↓  Table 2. Results of GSVA for GSE30219
 
logFCAveExprtP valueAdjusted P valueB
GSVA: gene set variation analysis.
CROSBY_E2F4_TARGETS0.9561080.0491595.7243142.421722828 × 10−86.513767243 × 10−78.781659
MONTERO_THYROID_CANCER_POOR_SURVIVAL_UP0.895730.0594865.8654181.133022404 × 10−84.452922465 × 10−79.503221
REACTOME_PHOSPHORYLATION_OF_EMI10.8843340.0238195.4510441.012865704 × 10−71.473852067 × 10−67.424824
FINETTI_BREAST_CANCER_KINOME_RED0.8835460.0424845.7464192.152025314 × 10−86.226890456 × 10−78.893763
KUMAMOTO_RESPONSE_TO_NUTLIN_3A_DN0.8788950.023745.3059022.118769427 × 10−72.438604173 × 10−66.726374
REACTOME_G2_M_DNA_REPLICATION_CHECKPOINT0.8622350.0225915.2668892.576893984 × 10−72.800822227 × 10−66.541305
FARMER_BREAST_CANCER_CLUSTER_20.8540060.0291596.0681453.713336651 × 10−92.403463232 × 10−710.56432
KEGG_MEDICUS_REFERENCE_RECRUITMENT_AND_FORMATION_OF_THE_MCC0.8507810.0102015.4968567.998335168 × 10−81.248156058 × 10−67.648502
REACTOME_UNWINDING_OF_DNA0.8446820.0491185.4583249.756630028 × 10−81.431236275 × 10−67.460265
REACTOME_TFAP2A_ACTS_AS_A_TRANSCRIPTIONAL_REPRESSOR_DURING_RETINOIC_ACID_INDUCED_CELL_DIFFERENTIATION0.8304870.0502755.5847295.063493258 × 10−89.553008036 × 10−78.081834
FERRARI_RESPONSE_TO_FENRETINIDE_DN−0.74442-0.02057−5.939377.567324611 × 10−93.576921019 × 10−79.886972
WP_NICOTINE_METABOLISM_IN_LIVER_CELLS−0.748730.015808−5.029218.287434703 × 10−76.414700134 × 10−65.438605
TONKS_TARGETS_OF_RUNX1_RUNX1T1_FUSION_SUSTAINED_IN_GRANULOCYTE_DN−0.762450.018083−5.089976.172673057 × 10−75.239527177 × 10−65.716414
REACTOME_ALTERNATIVE_COMPLEMENT_ACTIVATION−0.762950.038473−4.837032.066088753 × 10−61.273823858 × 10−54.57858
NAKAMURA_LUNG_CANCER_DIFFERENTIATION_MARKERS−0.76513−0.00215−6.528922.653051026 × 10−105.185639194 × 10−813.07993
KEGG_MEDICUS_REFERENCE_GLYCOGEN_DEGRADATION_AMYLASE−0.77418−0.00419−4.948431.220863401 × 10−68.489696264 × 10−65.073616
NAKAMURA_ALVEOLAR_EPITHELIUM−0.77422−0.01295−5.134274.970954288 × 10−74.522005209 × 10−65.920717
BLANCO_MELO_SARS_COV_1_INFECTION_MCR5_CELLS_DN−0.77506−0.03346−6.347967.605661673 × 10−109.335810772 × 10−812.07507
KEGG_MEDICUS_ENV_FACTOR_PHIP_TO_DNA_ADDUCTS−0.790250.027857−4.709153.735469016 × 10−62.058810541 × 10−54.022264
INAMURA_LUNG_CANCER_SCC_DN−0.89554−0.01957−7.030981.279898881 × 10−119.256228705 × 10−915.97746

 

↓  Table 3. Results of GSVA for GSE48000
 
logFCAveExprtP valueAdjusted P valueB
GSVA: gene set variation analysis.
JONES_TCOF1_TARGETS0.68535−0.008918.4190924.002037931 × 10−141.275010499 × 10−1221.61831
PALOMERO_GSI_SENSITIVITY_DN0.679906−0.016438.46763.039355649 × 10−141.017621299 × 10−1221.88801
WP_CELLULAR_PROTEOSTASIS0.672864−0.02437.4901067.009193031 × 10−121.224407826 × 10−1016.56116
KEGG_MEDICUS_REFERENCE_ELECTRON_TRANSFER_IN_COMPLEX_III0.6721060.0127978.4957152.590760166 × 10−148.845641371 × 10−1322.04454
KEGG_MEDICUS_ENV_FACTOR_ARSENIC_TO_ELECTRON_TRANSFER_IN_COMPLEX_IV0.650742−0.003388.9395152.042651044 × 10−159.783081028 × 10−1424.53563
KEGG_MEDICUS_REFERENCE_ELECTRON_TRANSFER_IN_COMPLEX_IV0.639247−0.004338.8148914.183298253 × 10−151.758930986 × 10−1323.83248
REACTOME_TRANSPORT_AND_SYNTHESIS_OF_PAPS0.6118060.0038447.8867667.926471046 × 10−131.780255857 × 10−1118.69348
KEGG_MEDICUS_VARIANT_MUTATION_CAUSED_ABERRANT_HTT_TO_ELECTRON_TRANSFER_IN_COMPLEX_III0.578126−0.00597.8702388.687263705 × 10−131.93311665 × 10−1118.60377
WP_DUAL_HIJACK_MODEL_OF_VIF_IN_HIV_INFECTION0.577045−0.041797.6775742.515284738 × 10−124.916361952 × 10−1117.56344
BIOCARTA_PARKIN_PATHWAY0.573573−0.013028.1562171.762596724 × 10−134.756380413 × 10−1220.16565
BIOCARTA_TERT_PATHWAY−0.655−0.00327−12.37023.189052496 × 10−243.843871276 × 10−2144.44821
REACTOME_NEF_AND_SIGNAL_TRANSDUCTION−0.655330.012557−9.535556.415944202 × 10−175.39536145 × 10−1527.93171
REACTOME_NEGATIVE_FEEDBACK_REGULATION_OF_MAPK_PATHWAY−0.655940.037132−9.148536.104456894 × 10−163.476175768 × 10−1425.72064
BIOCARTA_STAT3_PATHWAY−0.668190.001079−10.74435.123695773 × 10−201.323377423 × 10−1734.93634
WP_TCA_CYCLE_NUTRIENT_USE_AND_INVASIVENESS_OF_OVARIAN_CANCER−0.670750.033244−9.167825.458713298 × 10−163.235853653 × 10−1425.83035
KEGG_MEDICUS_PATHOGEN_KSHV_MIR1_2_TO_ANTIGEN_PROCESSING_AND_PRESENTATION_BY_MHC_CLASS_I_MOLECULES−0.673170.008885−7.701742.202486247 × 10−124.400105122 × 10−1117.69336
KEGG_MEDICUS_PATHOGEN_SHIGELLA_IPAB_C_D_TO_ITGA_B_TALIN_VINCULIN_SIGNALING_PATHWAY−0.68842−0.0191−11.71011.627713840 × 10−229.809688743 × 10−2040.58619
KEGG_MEDICUS_PATHOGEN_HTLV_1_P12_TO_JAK_STAT_SIGNALING_PATHWAY−0.706770.002033−10.85612.635139635 × 10−207.255917307 × 10−1835.5895
REACTOME_MATURATION_OF_SARS_COV_1_SPIKE_PROTEIN−0.71873−0.01869−9.429711.19034545 × 10−169.158062016 × 10−1527.32503
BIOCARTA_ION_PATHWAY−0.79573−0.00069−11.96363.592743199 × 10−232.886968758 × 10−2042.07005