Early and accurate cancer diagnosis remains one of the major challenges in modern healthcare, where diagnostic errors, limited interpretability of deep learning models, and the lack of integrated clinical reasoning can adversely affect patient outcomes. Although Convolutional Neural Networks (CNNs) have achieved remarkable performance in medical image analysis, they operate as black-box models and cannot effectively utilize clinical text or structured medical knowledge. This paper proposes a novel Self-Adaptive Hybrid CNN-LLM Framework that integrates visual, textual, and symbolic reasoning for cancer detection and clinical decision support. The proposed framework combines three complementary components: a deep CNN backbone for extracting spatial features from MRI, CT, and X-ray images; a biomedical Large Language Model (LLM) for interpreting clinical reports and generating diagnostic reasoning; and a Medical Knowledge Graph (MKG) based on the Unified Medical Language System (UMLS) for semantic validation of predictions. An attention-based transformer fusion module integrates information from the three modalities into a unified representation, while a Grad-CAM visualization module and an LLM-based explanation engine provide interpretable Explainable Artificial Intelligence (XAI) reports for clinicians. In addition, a self-adaptive reinforcement learning mechanism continuously refines the framework using clinician feedback without requiring complete model retraining. Experimental evaluation on three publicly available cancer datasets demonstrates that the proposed framework achieves 97.3% accuracy, 96.8% precision, 96.4% recall, and an AUC of 0.983, outperforming CNN-only and CNN+LLM baseline models by 4.2–7.9% across key performance metrics while providing substantially improved interpretability and clinical transparency. These results demonstrate that the proposed framework represents a reliable and trustworthy clinical decision support system that effectively bridges the gap between data-driven artificial intelligence and clinically meaningful medical reasoning.
Keywords
Convolutional Neural NetworksLarge Language ModelsCancer DetectionMedical Knowledge GraphSelf-Adaptive LearningMulti-Modal Fusion
References
H. Sung, J. Ferlay, R. L. Siegel, M. Laversanne, I. Soerjomataram, A. Jemal, and F. Bray, “Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries,” CA: A Cancer Journal for Clinicians, vol. 71, no. 3, pp. 209-249, 2021.
F. Mahmud et al., “An interpretable deep learning approach for skin cancer categorization,” in Proc. 26th International Conference on Computer and Information Technology (ICCIT), IEEE, 2023.
D. Ardila, A. P. Kiraly, S. Bharadwaj, B. Choi, J. J. Reicher, L. Peng, D. Tse, M. Etemadi, W. Ye, G. Corrado, D. P. Naidich, and S. Shetty, “End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography,” Nature Medicine, vol. 25, no. 6, pp. 954-961, 2019.
A. B. Arrieta, N. Díaz-Rodríguez, J. Del Ser, A. Bennetot, S. Tabik, A. Barbado, S. García, S. Gil-López, D. Molina, R. Benjamins, R. Chatila, and F. Herrera, “Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI,” Information Fusion, vol. 58, pp. 82-115, 2020.
K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, P. Payne, M. Seneviratne, P. Gamble, C. Kelly, N. Scharli, A. Chowdhery, P. Mansfield, B. Aguera-Arcas, D. Webster, G. S. Corrado, Y. Matias, K. Chou, J. Gottweis, N. Tomasev, Y. Liu, A. Rajkomar, J. Barral, C. Semturs, A. Karthikesalingam, and V. Natarajan, “Large language models encode clinical knowledge,” Nature, vol. 620, no. 7972, pp. 172-180, 2023.
R. Luo, L. Sun, Y. Xia, T. Qin, S. Zhang, H. Poon, and T.-Y. Liu, “BioGPT: Generative pre-trained transformer for biomedical text generation and mining,” Briefings in Bioinformatics, vol. 23, no. 6, art. bbac409, 2022.
O. Bodenreider, “The Unified Medical Language System (UMLS): Integrating biomedical terminology,” Nucleic Acids Research, vol. 32, no. Database issue, pp. D267-D270, 2004.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, Jun. 2016, pp. 770-778.
M. Tan and Q. V. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. 36th International Conference on Machine Learning (ICML), vol. 97, Long Beach, CA, USA, Jun. 2019, pp. 6105-6114.
G. Litjens, T. Kooi, B. E. Bejnordi, A. A. A. Setio, F. Ciompi, M. Ghafoorian, J. A. W. M. van der Laak, B. van Ginneken, and C. I. Sánchez, “A survey on deep learning in medical image analysis,” Medical Image Analysis, vol. 42, pp. 60-88, 2017.
J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. H. So, and J. Kang, “BioBERT: A pre-trained biomedical language representation model for biomedical text mining,” Bioinformatics, vol. 36, no. 4, pp. 1234-1240, 2020.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” International Journal of Computer Vision, vol. 128, no. 2, pp. 336-359, 2020.
R. M. Al-Amri, R. N. Hadi, and A. A. Mohammed, “Enhancement of the performance of machine learning algorithms to rival deep learning algorithms in predicting stock prices,” Babylonian Journal of Artificial Intelligence, vol. 2024, pp. 118-127, 2024.
N. Coudray, P. S. Ocampo, T. Sakellaropoulos, N. Narula, M. Snuderl, D. Fenyö, A. L. Moreira, N. Razavian, and A. Tsirigos, “Classification and mutation prediction from non-small cell lung cancer histopathology images using deep learning,” Nature Medicine, vol. 24, no. 10, pp. 1559-1567, 2018.
O. Tanimola et al., “Breast cancer classification using fine-tuned SWIN transformer model on mammographic images,” Analytics, vol. 3, no. 4, pp. 461-475, 2024.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, Long Beach, CA, USA, Dec. 2017, pp. 5998-6008.
P. Rajpurkar, E. Chen, O. Banerjee, and E. J. Topol, “AI in health and medicine,” Nature Medicine, vol. 28, no. 1, pp. 31-38, 2022.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Minneapolis, MN, USA, Jun. 2019, pp. 4171-4186.