Big Data Foundation Models for Personalized Medicine
Keywords:
Deep Learning; Foundation Models; Multi-task Learning; Genomics; Transcriptomics; Radiology; Clinical Medicine.Abstract
Deep learning models pretrained on multimodal Big Data may be used to address many problems in Precision Medicine. Choi et al. (2022) provided background on the core model architectures and training paradigms; on the prevalent types of multimodal medical data; and on the preliminary works aimed at combining model frameworks and training paradigm across existing modalities. Multimodal Foundation Models are introduced on the basis of learned medical representation and multimodal representation transfer and design. The explored areas are centered on data ecosystems that can support the future development of Deep Learning within Precision Medicine using multimodal Big Data.
As subtitled, the section on Data Ecosystems and Governance focuses on data acquisition, curation and quality assurance, privacy, security and ethical considerations. Applications involve the integration of genomics and transcriptomics; the use of medical image and radiomic data; and the building–testing of benchmark datasets to alleviate data bias. The addressed methodological challenges discuss data bias, fairness and generalizability; interpretability and clinician trust; benchmarking protocols; reproducibility; and open science practices.
References
[1] Alsentzer, E., Murphy, J. R., Boag, W., Weng, W.-H., Jin, D., Naumann, T., & McDermott, M. (2019). Publicly available clinical BERT embeddings. In Proceedings of the 2nd Clinical Natural Language Processing Workshop (pp. 72–78). Association for Computational Linguistics.
[2]Bannur, S., Boecking, B., Saha, O., Abu-Hanna, A., Schiratti, J.-B., & Shao, L. (2023). Learning to exploit temporal structure for biomedical vision-language processing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
[3]Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., … Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
[4]Chakraborty, S., Mondal, S., & Ghosh, S. (2020). A pre-trained biomedical language model for QA and IR. In Proceedings of the 28th International Conference on Computational Linguistics (COLING).
[5]Cui, H., Wang, C., Maan, H., Pang, J., Luo, F., Wang, B., & others. (2024). scGPT: Toward building a foundation model for single-cell biology. Nature Methods.
[6]Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT (pp. 4171–4186). Association for Computational Linguistics.
[7]Dou, Y., Lin, J., Zhang, Y., & others. (2025). MO-GCAN: Multi-omics integration based on graph convolutional attention networks. Bioinformatics, 41(8).
[8]Elnaggar, A., Heinzinger, M., Dallago, C., Rehawi, G., Wang, Y., Jones, L., … Rost, B. (2021). ProtTrans: Toward understanding the language of life through self-supervised learning. IEEE Transactions on Pattern Analysis and Machine Intelligence.
[9]Fu, X., Mo, S., Buendia, A., & others. (2025). A foundation model of transcription across human cell types. Nature, 637.
[10]Gu, Y., Tinn, R., Cheng, H., Lucas, M., Usuyama, N., Liu, X., Naumann, T., Gao, J., & Poon, H. (2021). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1).
[11]Huang, K., Altosaar, J., & Ranganath, R. (2020). ClinicalBERT: Modeling clinical notes and predicting hospital readmission. In Proceedings of the ACM Conference on Health, Inference, and Learning (CHIL).
[12]Ito, K., & others. (2025). Mouse-Geneformer: A deep learning model for mouse single-cell transcriptomics. Nature Communications.
[13]Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., … Hassabis, D. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596, 583–589.
[14]Khan, S. A., & others. (2023). Learning the transcriptional grammar in single-cell RNA sequencing. Nature Machine Intelligence, 5.
[15]Kipf, T. N., & Welling, M. (2017). Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations (ICLR).
[16]Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., & Kang, J. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240.
[17]Li, Y., Rao, S., Ayala Solares, J. R., Hassaine, A., Ramakrishnan, R., Canoy, D., Zhu, Y., Rahimi, K., & Salimi-Khorshidi, G. (2020). BEHRT: Transformer for electronic health records. Scientific Reports, 10, 7155.
[18]Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., … Rives, A. (2023). Evolutionary-scale prediction of atomic-level protein structure with a language model. Science, 379(6637), 1123–1130.
[19]Lu, Y., & others. (2025). Integrating language into medical visual recognition and reasoning: A review of medical vision-language models. Medical Image Analysis, 99, 103xxx.
[20]Luo, R., Sun, L., Xia, Y., Qin, T., Zhang, S., Poon, H., & Liu, T.-Y. (2022). BioGPT: Generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23(6), bbac409.
[21]Malik, V., & others. (2021). Deep learning assisted multi-omics integration for survival and drug response prediction in breast cancer. Frontiers in Genetics, 12, 123456.
[22]Ouyang, D., & others. (2023). Integration of multi-omics data using adaptive graph learning for disease classification prediction. Computers in Biology and Medicine, 164, 107xxx.
[23]Park, S., & others. (2025). Advancing protein structure prediction beyond AlphaFold2. Current Opinion in Structural Biology, 82, 1026xx.
[24]Peng, C., Yang, X., Chen, A., Smith, K. E., & others. (2023). A study of generative large language model for medical research and healthcare. npj Digital Medicine, 6, 210.
[25]Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., … Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML).
[26]Rasmy, L., Xiang, Y., Xie, Z., Tao, C., & Zhi, D. (2021). Med-BERT: Pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. npj Digital Medicine, 4, 86.
[27]Shin, H.-C., Zhang, Y., Bakhturina, E., Puri, R., Patwary, M., Shoeybi, M., & Mani, R. (2020). BioMegatron: Larger biomedical domain language model. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 4700–4706). Association for Computational Linguistics.
[28]Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., … Natarajan, V. (2023). Large language models encode clinical knowledge. Nature, 620, 172–180.
[29]Singhal, K., Tu, T., Gottweis, J., & others. (2025). Toward expert-level medical question answering with large language models. Nature Medicine, 31.
[30]Tejada-Lapuerta, A., & others. (2025). Nicheformer: A foundation model for single-cell and spatial transcriptomics. Nature Methods, 22.
[31]Theodoris, C. V., Xiao, L., Chopra, A., Chaffin, M., & others. (2023). Transfer learning enables predictions in network biology. Nature, 618, 616–624.
[32]Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, 30.
[33]Wang, T., Shao, W., Huang, Z., Tang, H., Zhang, J., Ding, Z., & Huang, K. (2021). MOGONET integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nature Communications, 12, 3445.
[34]Wang, W., Yang, F., Fang, Y., & others. (2022). scBERT as a large-scale pretrained deep language model for cell type annotation of single-cell RNA-seq data. Nature Machine Intelligence, 4.
[35]Wang, Z., Wu, Z., Agarwal, D., & Sun, J. (2022). MedCLIP: Contrastive learning from unpaired medical images and text. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 3876–3887). Association for Computational Linguistics.
[36]Wissel, D., & others. (2023). Systematic comparison of multi-omics survival models on pan-cancer datasets. Patterns, 4(6), 1007xx.
[37]Xian, S., & others. (2025). Transformer patient embedding using electronic health records for phenotyping and longitudinal analysis. npj Digital Medicine, 8.
[38]Yang, X., Chen, A., PourNejatian, N., Shin, H. C., Smith, K. E., Parisien, C., … Wu, Y. (2022). A large language model for electronic health records. npj Digital Medicine, 5, 194.
[39]Yang, Z., & others. (2023). TransformEHR: Transformer-based encoder-decoder modeling of longitudinal electronic health records. Nature Communications, 14, 7xxx.
[40]Zhang, X., & others. (2023). Knowledge-enhanced visual-language pre-training on chest radiographs for automated diagnosis. Nature Communications, 14, 4xxx.
[41]Chen, I. Y., Pierson, E., Rose, S., Joshi, S., Ferryman, K., & Ghassemi, M. (2021). Ethical machine learning in healthcare. Annual Review of Biomedical Data Science, 4, 123–144.
[42]Esteva, A., Robicquet, A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., & Dean, J. (2019). A guide to deep learning in healthcare. Nature Medicine, 25, 24–29.
[43]Rajkomar, A., Dean, J., & Kohane, I. (2019). Machine learning in medicine. New England Journal of Medicine, 380(14), 1347–1358.
[44]Beam, A. L., & Kohane, I. S. (2018). Big data and machine learning in health care. JAMA, 319(13), 1317–1318.
[45]Miotto, R., Wang, F., Wang, S., Jiang, X., & Dudley, J. T. (2018). Deep learning for healthcare: Review, opportunities and challenges. Briefings in Bioinformatics, 19(6), 1236–1246.
[46]Shickel, B., Tighe, P. J., Bihorac, A., & Rashidi, P. (2018). Deep EHR: A survey of recent advances in deep learning techniques for electronic health record analysis. IEEE Journal of Biomedical and Health Informatics, 22(5), 1589–1604.
[47]Chen, R. J., Lu, M. Y., Chen, T. Y., Williamson, D. F. K., & Mahmood, F. (2022). Synthetic data in machine learning for medicine and healthcare. Nature Biomedical Engineering, 6, 106–118.
[4]8]Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25, 44–56.
[49]Lotfollahi, M., Naghipourfar, M., Theis, F. J., & Wolf, F. A. (2022). Deep learning for single-cell analysis. Nature Reviews Genetics, 23, 781–799.
[50]He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 770–778).
[51]Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI) (pp. 234–241). Springer.
[52]Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of CVPR (pp. 4700–4708).
[53]Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020). A simple framework for contrastive learning of visual representations. In Proceedings of ICML.
[54]He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (2020). Momentum contrast for unsupervised visual representation learning. In Proceedings of CVPR.
[55]Kather, J. N., & others. (2019). Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer. Nature Medicine, 25, 1054–1056.
[56]Echle, A., & others. (2020). Deep learning in cancer pathology: A new generation of clinical biomarkers. British Journal of Cancer, 124, 686–696.
[57]Campanella, G., Hanna, M. G., Geneslaw, L., Miraflor, A., Werneck Krauss Silva, V., Busam, K. J., … Fuchs, T. J. (2019). Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nature Medicine, 25, 1301–1309.
[58]Lu, M. Y., Chen, R. J., Kong, D., Lipkova, J., Singh, R., Williamson, D. F. K., & Mahmood, F. (2021). Federated learning for computational pathology on gigapixel whole slide images. Medical Image Analysis, 76, 102298.
[59]Johnson, A. E. W., Pollard, T. J., Shen, L., Lehman, L. W. H., Feng, M., Ghassemi, M., … Mark, R. G. (2016). MIMIC-III, a freely accessible critical care database. Scientific Data, 3, 160035.
[60]Wu, Y., & others. (2023). Foundation models in healthcare: Opportunities, risks, and pathways to safe deployment. Nature Medicine, 29.
[61]Boecking, B., & others. (2022). Making the most of text semantics to improve biomedical vision-language processing. arXiv.
[62]Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3), e0118432.
[63]Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems, 30.
[64]Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?” Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1135–1144). ACM.
[65]Allen-Zhu, Z., Li, Y., & Song, Z. (2019). A convergence theory for deep learning via over-parameterization. In Proceedings of ICML.
[66]Dwork, C., Hardt, M., Pitassi, T., Reingold, O., & Zemel, R. (2012). Fairness through awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference (pp. 214–226). ACM.
[67]Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.
[68]Chen, J. H., & Asch, S. M. (2017). Machine learning and prediction in medicine—Beyond the peak of inflated expectations. New England Journal of Medicine, 376(26), 2507–2509.
[69]Price, W. N., & Cohen, I. G. (2019). Privacy in the age of medical big data. Nature Medicine, 25, 37–43.
[70]Park, Y., & others. (2024). Benchmarking and auditing large language models for clinical safety and reliability. npj Digital Medicine, 7.
Additional Files
Published
Data Availability Statement
None