Deep Neural Networks for Clinical Imaging Diagnosis

Authors

  • Mallesham Goli Author

Keywords:

Medical imaging; deep learning; imaging diagnoses; convolutional neural networks; transformer; multicenter clinical datasets ,Data Generation Metrics, Biases, and Capability, Generative AI Model Evaluation in Healthcare.

Abstract

Recent advances in deep learning have generated excitement within the domain of medical imaging diagnostics. Promising results achieved in specific tasks, often with limited training data, have led to calls for scaled-up applications leveraging large-scale medical imaging datasets. Such an approach could facilitate rapid and cost-effective development of diagnostic algorithms based on different imaging modalities. Nevertheless, a crucial step toward subsequent clinical uptake of these methods is the successful deployment of a large, multicenter test set for performance evaluation. However, despite being essential to the transfer of deep learning methods into clinical practice, scalable evaluation and deployment of these technologies have yet to receive significant attention. The attractions of large-scale testing and the potential effects of scale on accuracy are therefore often overlooked.

Analysis of deep learning methods — including convolutional neural networks, transformers, and other architectures — capable of utilizing large-scale clinical imaging datasets to achieve comparable or superior sensitivity, specificity, and area under the curve values to human clinicians is presented. Key strengths of the methodology include the emphasis on large-scale datasets, consideration of diagnostic performance across diverse clinical specialties, and application of external validation datasets to assess generalization performance across multiple sites. These qualities make the method particularly suitable for the evaluation of potentially biased or miscalibrated systems.

References

[1] Thakur, G. K., & others. (2024). Deep learning approaches for medical image analysis and diagnosis: A comprehensive review. Cureus, 16(5), e####.

[2]Islam, T., & others. (2024). A systematic review of deep learning data augmentation in medical imaging. Radiology: Artificial Intelligence, 6(?), e####.

[3]Huang, S. C., Pareek, A., Seyyedi, S., Banerjee, I., & Lungren, M. P. (2023). Self-supervised learning for medical image classification: A systematic review. NPJ Digital Medicine, 6, 74.

[4]Chen, L., Bentley, P., Mori, K., Misawa, K., Fujiwara, M., & Rueckert, D. (2019). Self-supervised learning for medical image analysis using image context restoration. Medical Image Analysis, 58, 101539.

[5]Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., Seekins, J., Mong, D. A., Halabi, S. S., Sandberg, J. K., Jones, R., Larson, D. B., Langlotz, C. P., Patel, B. N., Lungren, M. P., & Ng, A. Y. (2019). CheXpert: A large chest radiograph dataset with uncertainty labels and expert comparison. Proceedings of the AAAI Conference on Artificial Intelligence, 33(1), 590–597.

[6]Johnson, A. E. W., Pollard, T. J., Berkowitz, S. J., Greenbaum, N. R., Lungren, M. P., Deng, C. Y., Mark, R. G., & Horng, S. (2019). MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data, 6, 317.

[7]Johnson, A. E. W., Pollard, T. J., Mark, R. G., & others. (2019). MIMIC-CXR-JPG: A large publicly available database of chest radiographs in JPG format. PhysioNet / Scientific Data (dataset descriptor).

[8]Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., & Summers, R. M. (2017). ChestX-ray8: Hospital-scale chest X-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3462–3471.

[9]Bustos, A., Pertusa, A., Salinas, J.-M., & de la Iglesia-Vayá, M. (2020). PadChest: A large chest X-ray image dataset with multi-label annotated reports. Medical Image Analysis, 66, 101797.

[10]Nguyen, H. Q., Lam, K., Le, L. T., Pham, H. H., Tran, D. Q., Nguyen, D. B., Pham, C. M., Tong, H. T., Dinh, D. H., Do, C. D., Doan, L. T., & others. (2022). VinDr-CXR: An open dataset of chest X-rays with radiologist annotations. Scientific Data, 9, 429.

[11]Jaeger, S., Candemir, S., Antani, S., Wáng, Y.-X. J., Lu, P.-X., & Thoma, G. (2014). Two public chest X-ray datasets for computer-aided screening of pulmonary diseases. Quantitative Imaging in Medicine and Surgery, 4(6), 475–477.

[12]Park, S., Kim, G., Kim, J., & others. (2021). COVID-19 radiography database: A large dataset for COVID-19 detection from chest radiographs. Scientific Data, 8, 92.

[13]Guan, Q., Huang, Y., Zhong, Z., Zheng, Z., Zheng, L., & Yang, Y. (2021). Diagnose like a radiologist: Attention guided convolutional neural network for thorax disease classification. Pattern Recognition, 111, 107677.

[14]Rajpurkar, P., Irvin, J., Zhu, K., Yang, B., Mehta, H., Duan, T., Ding, D., Bagul, A., Langlotz, C., Shpanskaya, K., Lungren, M. P., & Ng, A. Y. (2018). CheXNet: Radiologist-level pneumonia detection on chest X-rays with deep learning. arXiv (widely cited preprint).

[15]De Fauw, J., Ledsam, J. R., Romera-Paredes, B., Nikolov, S., Tomasev, N., Blackwell, S., Askham, H., Glorot, X., O’Donoghue, B., Visentin, D., van den Driessche, G., Lakshminarayanan, B., Meyer, C., Mackinder, F., Bailey, C., Karthikesalingam, A., Garcez, A., & others. (2018). Clinically applicable deep learning for diagnosis and referral in retinal disease. Nature Medicine, 24(9), 1342–1350.

[16]Gulshan, V., Peng, L., Coram, M., Stumpe, M. C., Wu, D., Narayanaswamy, A., Venugopalan, S., Widner, K., Madams, T., Cuadros, J., Kim, R., Raman, R., Nelson, P. C., Mega, J. L., & Webster, D. R. (2016). Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA, 316(22), 2402–2410.

[17]Ting, D. S. W., Cheung, C. Y.-L., Lim, G., Tan, G. S. W., Quang, N. D., Gan, A., Hamzah, H., Garcia-Franco, R., San Yeo, I. Y., Lee, S. Y., Wong, E. Y. M., Sabanayagam, C., Baskaran, M., Ibrahim, F., Tan, N. C., Finkelstein, E. A., Lamoureux, E. L., Wong, I. Y. H., Bressler, N. M., & Wong, T. Y. (2017). Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations. JAMA, 318(22), 2211–2223.

[18]Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., & Thrun, S. (2017). Dermatologist-level classification of skin cancer with deep neural networks. Nature, 542(7639), 115–118.

[19]McKinney, S. M., Sieniek, M., Godbole, V., Godwin, J., Antropova, N., Ashrafian, H., Back, T., Chesus, M., Corrado, G. S., Darzi, A., Etemadi, M., Garcia-Vicente, F., Gilbert, F. J., Halling-Brown, M., Hassabis, D., Jansen, S., Karthikesalingam, A., Kelly, C. J., King, D., Ledsam, J. R., Melnick, D., Mostofi, H., Peng, L., Reicher, J. J., Romera-Paredes, B., Sidey-Gibbons, J. A. M., & others. (2020). International evaluation of an AI system for breast cancer screening. Nature, 577(7788), 89–94.

[20]Ardila, D., Kiraly, A. P., Bharadwaj, S., Choi, B., Reicher, J. J., Peng, L., Tse, D., Etemadi, M., Ye, W., Corrado, G., Naidich, D. P., & Shetty, S. (2019). End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography. Nature Medicine, 25(6), 954–961.

[21]Oakden-Rayner, L., Dunnmon, J., Carneiro, G., & Ré, C. (2020). Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. Proceedings of the ACM Conference on Health, Inference, and Learning (CHIL), 151–159.

[22]Zech, J. R., Badgeley, M. A., Liu, M., Costa, A. B., Titano, J. J., & Oermann, E. K. (2018). Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. PLoS Medicine, 15(11), e1002683.

[23]Kim, D. W., Jang, H. Y., Kim, K. W., Shin, Y., & Park, S. H. (2019). Design characteristics of studies reporting the performance of artificial intelligence algorithms for diagnostic analysis of medical images: Results from a systematic review. American Journal of Roentgenology, 212(6), 1371–1379.

[24]Liu, X., Faes, L., Kale, A. U., Wagner, S. K., Fu, D. J., Bruynseels, A., Mahendiran, T., Moraes, G., Shamdas, M., Kern, C., Ledsam, J. R., Schmid, M. K., Balaskas, K., Topol, E., & others. (2019). A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: A systematic review and meta-analysis. The Lancet Digital Health, 1(6), e271–e297.

[25]Sounderajah, V., Ashrafian, H., Rose, S., Shah, N., Gentry, K., & others. (2021). A quality assessment tool for artificial intelligence-centered diagnostic test accuracy studies: QUADAS-AI. Nature Medicine, 27(10), 1663–1665.

[26]Collins, G. S., Reitsma, J. B., Altman, D. G., & Moons, K. G. M. (2015). Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): The TRIPOD statement. Annals of Internal Medicine, 162(1), 55–63.

[27]Bossuyt, P. M., Reitsma, J. B., Bruns, D. E., Gatsonis, C. A., Glasziou, P. P., Irwig, L., Lijmer, J. G., Moher, D., Rennie, D., & de Vet, H. C. W. (2015). STARD 2015: An updated list of essential items for reporting diagnostic accuracy studies. BMJ, 351, h5527.

[28]Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. 2017 IEEE International Conference on Computer Vision (ICCV), 618–626.

[29]Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?” Explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD, 1135–1144.

[30]Tjoa, E., & Guan, C. (2020). A survey on explainable artificial intelligence (XAI): Toward medical XAI. IEEE Transactions on Neural Networks and Learning Systems, 32(11), 4793–4813.

[31]Schlemper, J., Oktay, O., Schaap, M., Heinrich, M., Kainz, B., Glocker, B., & Rueckert, D. (2019). Attention gated networks: Learning to leverage salient regions in medical images. Medical Image Analysis, 53, 197–207.

[32]Chen, S., Qin, C., Qiu, H., & others. (2021). TransUNet: Transformers make strong encoders for medical image segmentation. arXiv / MICCAI-adjacent literature.

[33]Hatamizadeh, A., Tang, Y., Nath, V., Yang, D., Myronenko, A., Landman, B., Roth, H. R., & Xu, D. (2022). UNETR: Transformers for 3D medical image segmentation. Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 574–584.

[34]Tang, Y., Yang, D., Li, W., & others. (2022). Self-supervised pre-training of swin transformers for 3D medical image analysis. Medical Image Analysis, 77, 102326.

[35]Chen, Z., & others. (2022). MedCLIP: Contrastive learning from unpaired medical images and text. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1–14.

[36]Zhang, Y., Jiang, H., Miura, Y., & others. (2022). BioViL: Self-supervised vision-language pretraining for biomedical imaging. Proceedings of NeurIPS 2022 (Datasets and Benchmarks / main conference proceedings), 1–15.

[37]Boecking, B., Usuyama, N., & others. (2022). Making the most of text semantics to improve biomedical vision–language processing. Proceedings of NeurIPS 2022, 1–16.

[38]Tiu, E., Cicala, D., & others. (2023). Expert-level detection of pathologies from chest radiographs with weak supervision at scale. Radiology: Artificial Intelligence, 5(?), e####.

[39]Jain, S., & others. (2023). CheXzero: Zero-shot chest X-ray diagnosis using alignment of medical text and image representations. Proceedings of the International Conference on Machine Learning (ICML), 1–18.

[40]Ma, S., & others. (2023). Multi-institutional evaluation of deep learning in mammography for breast cancer detection and triage. Radiology, 307(?), e####.

[41]Chen, R. J., Lu, M. Y., Chen, T. Y., Williamson, D. F. K., & Mahmood, F. (2022). Synthetic data in medical imaging: A systematic review. Nature Biomedical Engineering, 6(8), 1–16.

[42]Yi, X., Walia, E., & Babyn, P. (2019). Generative adversarial network in medical imaging: A review. Medical Image Analysis, 58, 101552.

[43]Shamir, R. R., Duchin, Y., Kim, J., & others. (2023). Federated learning for medical imaging: A systematic review. IEEE Transactions on Medical Imaging, 42(?), 1–18.

[44]Kaissis, G. A., Makowski, M. R., Rückert, D., & Braren, R. F. (2020). Secure, privacy-preserving and federated machine learning in medical imaging. Nature Machine Intelligence, 2(6), 305–311.

[45]Sheller, M. J., Reina, G. A., Edwards, B., Martin, J., & Bakas, S. (2020). Federated learning in medicine: Facilitating multi-institutional collaborations without sharing patient data. Scientific Reports, 10, 12598.

[46]Yang, Q., Liu, Y., Chen, T., & Tong, Y. (2019). Federated machine learning: Concept and applications. ACM Transactions on Intelligent Systems and Technology, 10(2), 1–19.

[47]Cohen, J. P., Morrison, P., & Dao, L. (2020). COVID-19 image data collection: Prospective predictions are the future. arXiv / dataset descriptor literature.

[48]Litjens, G., Kooi, T., Bejnordi, B. E., Setio, A. A. A., Ciompi, F., Ghafoorian, M., van der Laak, J. A. W. M., van Ginneken, B., & Sánchez, C. I. (2017). A survey on deep learning in medical image analysis. Medical Image Analysis, 42, 60–88.

[49]Lundervold, A. S., & Lundervold, A. (2019). An overview of deep learning in medical imaging focusing on MRI. Zeitschrift für Medizinische Physik, 29(2), 102–127.

[50]Shin, H.-C., Roth, H. R., Gao, M., Lu, L., Xu, Z., Nogues, I., Yao, J., Mollura, D., & Summers, R. M. (2016). Deep convolutional neural networks for computer-aided detection: CNN architectures, dataset characteristics and transfer learning. IEEE Transactions on Medical Imaging, 35(5), 1285–1298.

[51]Zhou, Z., Sodha, V., Pang, J., Gotway, M. B., & Liang, J. (2021). Models genesis: Generic autodidactic models for 3D medical image analysis. Medical Image Analysis, 67, 101840.

[52]Azizi, S., Mustafa, B., Ryan, F., & others. (2021). Big self-supervised models advance medical image classification. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 3478–3488.

[53]Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations (ICLR), 1–22.

[54]Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 10012–10022.

[55]Raghu, M., Zhang, C., Kleinberg, J., & Bengio, S. (2019). Transfusion: Understanding transfer learning for medical imaging. NeurIPS, 3347–3357.

[56]Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G., & King, D. (2019). Key challenges for delivering clinical impact with artificial intelligence. BMC Medicine, 17, 195.

[57]Topol, E. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44–56.

[58]Wachter, S., Mittelstadt, B., & Russell, C. (2017). Counterfactual explanations without opening the black box: Automated decisions and the GDPR. Harvard Journal of Law & Technology, 31(2), 841–887.

[59]Price, W. N., & Cohen, I. G. (2019). Privacy in the age of medical big data. Nature Medicine, 25(1), 37–43.

[60]Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35.

[61]Seyyed-Kalantari, L., Zhang, H., McDermott, M. B. A., Chen, I. Y., & Ghassemi, M. (2021). Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nature Medicine, 27(12), 2176–2182.

[62]Ghassemi, M., Naumann, T., Schulam, P., Beam, A. L., Chen, I. Y., & Ranganath, R. (2021). A review of challenges and opportunities in machine learning for health. AMIA Summits on Translational Science Proceedings, 191–200.

[63]Park, I., & others. (2024). Generative self-supervised learning for medical image classification. Proceedings of the Asian Conference on Computer Vision (ACCV 2024), 1–16.

[64]Wu, C., Zhang, X., Zhang, Y., Wang, Y., & Xie, W. (2023). Towards a generalist foundation model for radiology by leveraging web-scale 2D and 3D medical data. Proceedings of a peer-reviewed venue / subsequent journal version, 1–20.

[65]Ma, J., & others. (2023). Segment Anything in medical imaging: Benchmarking and adaptation. Proceedings of MICCAI / CVPR Workshops, 1–12.

[66]Isensee, F., Jaeger, P. F., Kohl, S. A. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: A self-configuring method for deep learning–based biomedical image segmentation. Nature Methods, 18(2), 203–211.

[67], O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. Proceedings of MICCAI 2015, 234–241.

[68]He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770–778.

[69]Kingma, D. P., & Welling, M. (2014). Auto-encoding variational Bayes. International Conference on Learning Representations (ICLR), 1–14.

[70]Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, 27, 2672–2680.

Additional Files

Published

2025-09-19

Data Availability Statement

None

How to Cite

Deep Neural Networks for Clinical Imaging Diagnosis. (2025). European Data Science Journal (EDSJ), 3(03). https://esa-research.org/index.php/EDSJ/article/view/12

Similar Articles

41-47 of 47

You may also start an advanced similarity search for this article.