Secure Cloud-Native Architectures for Generative AI in Enterprise Workflows

Authors

  • Yannis Ioannidis Author

Keywords:

Cloud-Native AI Infrastructure,Secure Generative AI Deployment,Enterprise Workflow Automation,Deep Learning Architecture Design,Kubernetes-Based AI Orchestration,AI Model Security and Governance,Scalable Generative AI Systems,MLOps for Enterprise AI Platforms,Zero-Trust AI Deployment Frameworks,Containerized Deep Learning Pipelines.

Abstract

Generative artificial intelligence (GenAI) is a rapidly developing technology with potential for multiple applications, yet it is complex, resource-intensive, and prone to risks. Deployment of GenAI in enterprise workflow platforms requires approaches that enable secure operation while maintaining availability, reliability, and quality. A synthesis of cloud-native architectural patterns with contemporary risk frameworks provides insights into essential security aspects for GenAI. Findings indicate that careful consideration of the safeguards available for prompt injection mitigation, model inversion protection, and data privacy when developing GenAI within a service-mesh architecture can reduce the likelihood of future attack success or damage.

Cloud-native generative artificial intelligence (GenAI) deployment in enterprise workflow platforms is increasingly common, especially for support documentation creation. However, safeguarding the system against attacks that target the availability, reliability, or data privacy of the service remains challenging. Leveraging cloud-native patterns of scalability, resilience, composability, portability, and observability can guide security-enhancing measures. Mapping the security-by-design concept to GenAI services within a service-mesh architecture identifies a range of security controls rooted in established identity and access management principles, the concept of security through obscurity, the defense-in-depth principle, and the auditing of logs and monitoring alerts.

References

1. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., ... Liang, P. (2022). On the opportunities and risks of foundation models. arXiv.

2. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., ... Amodei, D. (2022). Language models are few-shot learners. Communications of the ACM, 65(4), 126–135.

3. Radha, S., Gottimukkala, V. R. R., Thottara, S., Vandhana, K., & J, Gokulraj. (2025). Adaptive Video Streaming Over 5G Networks Using Deep Reinforcement Learning with Closed-Loop Feedback Mechanism for Bitrate Control. In 2025 International Conference on Communication, Computer, and Information Technology (IC3IT) (pp. 1–6). IEEE. 2025 International Conference on Communication, Computer, and Information Technology (IC3IT). https://doi.org/10.1109/ic3it66137.2025.11341184

4. Chiang, W. L., Zheng, L., Sheng, Y., Li, T., Zhuang, Z., Wu, Y., ... Stoica, I. (2023). Vicuna: An open-source chatbot impressing GPT-4 with 90% ChatGPT quality. arXiv.

5. Ding, Y., Qin, X., Wei, Z., & Choi, B. Y. (2023). Cloud-native computing: Architecture, security, and applications. IEEE Access, 11, 109215–109236.

6. Kummari, D. N., Burugulla, J. K. R., Malempati, M., Amistapuram, K., Garapati, R. S., & Nagabhyru, K. C. (2025, December). Enhancing Audit Compliance and Operational Efficiency in Manufacturing and Commercial Insurance Through Agentic AI and Data Engineering Frameworks. In 2025 IEEE International Conference on Communication Networks and Computing (CNC) (pp. 714-720). IEEE.

7. Dong, X., Li, Y., Sun, H., Zhang, C., & Chen, J. (2024). Secure deployment architectures for generative AI services in cloud environments. Future Internet, 16(4), 112.

8. Guo, Y., Li, X., Wang, H., & Zhang, Y. (2024). Zero-trust security architecture for cloud-native AI platforms. Future Internet, 16(5), 154.

9. He, Y., Bian, S., Chen, L., Hui, Y., Lentz, M., Li, B., ... Zhuo, D. (2024). Computing in the era of large generative models: From cloud-native to AI-native. arXiv.

10. Kolla, T. (2025). Generative AI for Intelligent Medical Coding and Healthcare Analytics. International Journal of Advanced Research in Computer Science & Technology (IJARCST), 8(6), 13285-13299.

11. Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D., ... Sayed, W. (2024). Mixtral of experts. arXiv.

12. Kandpal, N., Deng, H., Roberts, A., Wallace, E., & McCallum, A. (2023). Large language models struggle to learn long-tail knowledge. Proceedings of the 40th International Conference on Machine Learning.

13. Inala, R., Kaulwar, P. K., Nagabhyru, K. C., Adusupalli, B., & Arun Raj, S. R. (2025, October). Leveraging IEC 61850 for Interoperable and Resilient Smart Grid Communication Architecture. In International Conference on Microelectronics, Electromagnetics and Telecommunication (pp. 549-566). Cham: Springer Nature Switzerland.

14. Karpathy, A. (2023). Software 2.0 and the rise of AI-native application development. arXiv.

15. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., ... Riedel, S. (2022). Retrieval-augmented generation for knowledge-intensive NLP tasks. Journal of Machine Learning Research, 23, 1–48.

16. Amistapuram, K., Pandiri, L., Raju, V. R., Paleti, S., Singireddy, S., & Sheelam, G. K. (2025). AI-Based Cloud Infrastructure and MLOps Frameworks for Scalable Data Engineering Across Banking and Insurance. In 2025 IEEE International Conference on Communication Networks and Computing (CNC) (pp. 186–192). IEEE. 2025 IEEE International Conference on Communication Networks and Computing (CNC). https://doi.org/10.1109/cnc68716.2025.11484532

17. Li, X., Zhang, Y., Wang, Z., & Chen, H. (2023). Kubernetes-based orchestration for scalable AI microservices. IEEE Access, 11, 88431–88447.

18. Lu, Y., Bian, S., Chen, L., He, Y., Hui, Y., Li, B., ... Zhuo, D. (2024). AI-native computing: Challenges and opportunities for cloud infrastructure. arXiv.

19. OpenAI. (2023). GPT-4 technical report. arXiv.

20. Sudhakar, A. V. V., Inala, R., Verma, A. K., Nag, K., Pandey, V., & Anand, P. S. (2025, September). Hybrid Rule-Based and Machine Learning Framework for Embedding Anti-Discrimination Law in Automated Decision Systems. In 2025 International Conference on Intelligent Communication Networks and Computational Techniques (ICICNCT) (pp. 1-6). IEEE.

21. Patel, D., Raut, G., Cheetirala, S. N., Nadkarni, G. N., Freeman, R., Glicksberg, B. S., ... Klang, E. (2024). Cloud platforms for developing generative AI solutions: A scoping review of tools and services. arXiv.

22. Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., ... Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv.

23. Wang, Z., Chen, H., Li, X., & Zhang, Y. (2023). Secure microservice architectures for enterprise cloud applications. IEEE Access, 11, 97624–97641.

24. Loganathan, R. (2025). AGENTIC AI FRAMEWORKS FOR AUTONOMOUS RISK DETECTION AND COMPLIANCE REMEDIATION IN ENTERPRISE DATA CENTER OPERATIONS. Lex Localis-Journal of Local Self-Government, 23 (S6), 9672–9697.

25. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., ... Zhou, D. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research.

26. Zhang, C., Liu, J., Wang, H., & Sun, X. (2024). Secure cloud-native deployment framework for enterprise generative AI applications. Journal of Cloud Computing, 13(1), 58.

27. Nigam, N., Sireesha, B., Ediga, P., Segireddy, A. R., & Bokde, S. (2025, December). Comparative Evaluation of Cloud Security Algorithms Using Multiple Classifiers with an Optimized Intrusion Detection System. In 2025 IEEE 5th International Conference on ICT in Business Industry & Government (ICTBIG) (pp. 1-6). IEEE.

28. Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., ... Wen, J. R. (2023). A survey of large language models. arXiv.

29. Deng, S., Zhao, H., Huang, B., Zhang, C., Chen, F., Deng, Y., Yin, J., Dustdar, S., & Zomaya, A. Y. (2024). Cloud-native computing: A survey from the perspective of services. Proceedings of the IEEE, 112(1), 12–46.

30. Marin, E., Perino, D., & Di Pietro, R. (2022). Serverless computing: A security perspective. Journal of Cloud Computing, 11, Article 69.

31. Amistapuram¹, K., Kolla, T., Bandi, V. D. V. K., Kolla⁴, S. K., & Rani, P. S. Journal of Rare Cardiovascular Diseases.

32. OpenAI. (2023). GPT-4 technical report. arXiv:2303.08774.

33. Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., et al. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288.

34. Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., et al. (2023). LLaMA: Open and efficient foundation language models. arXiv:2302.13971.

35. Singh, H., Bose, D., Nagubandi, A. R., Prabhu, S., & Naik, S. G. Cryptocurrency Market Spillovers: Risk Contagion Across Global Financial Systems.

36. Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., et al. (2023). A survey of large language models. arXiv.

37. Lu, Y., Bian, S., Chen, L., He, Y., Hui, Y., Lentz, M., Li, B., et al. (2024). Computing in the era of large generative models: From cloud-native to AI-native. arXiv:2401.12230.

38. Patel, D., Raut, G., Cheetirala, S. N., Nadkarni, G. N., Freeman, R., Glicksberg, B. S., Klang, E., & Timsina, P. (2024). Cloud platforms for developing generative AI solutions: A scoping review of tools and services. arXiv:2412.06044.

39. Rani, P. S., Kummari, D. N., Yellanki, S. K., Meda, R., Koppolu, H. K. R., & Inala, R. (2025, July). Blockchain and AI for Securing Electrical Infrastructure. In 2025 2nd International Conference on Computing and Data Science (ICCDS) (pp. 1-6). IEEE.

40. Ramesh, G., Pai, T. V., Birau, R., Poojary, K. K., Abhay, Shingad, A. R., Sowjanya, N., Popescu, V., Mitroi, A. T., Nioata Chireac, R. M., & Raj, K. M. K. (2025). A comprehensive review on scaling machine learning workflows using cloud technologies and DevOps. IEEE Access.

41. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., et al. (2022). On the opportunities and risks of foundation models. arXiv.

42. Mangalampalli, B. M., Bandi, V. D. V. K., Kolla, S. K., & Kumar, M. V. K. (2025). Towards Self-Evolving Healthcare Intelligence: Integrating Advanced Learning Systems with Real-Time Clinical Data Pipelines. Cultura: International Journal of Philosophy of Culture and Axiology, 22(12s), 464-486.

43. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., et al. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research.

44. Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D., et al. (2024). Mixtral of experts. arXiv.

45. Chiang, W.-L., Zheng, L., Sheng, Y., Li, T., Zhuang, Z., Wu, Y., et al. (2023). Vicuna: An open-source chatbot impressing GPT-4 with 90% ChatGPT quality. arXiv.

46. Kumar, I., Nagabhyru, K. C., IG, N., MV, P., & KV, S. (2025, October). Adaptive Meta-Knowledge Transfer Network with Feature Hallucination and Attention for Low-Shot Object Detection in Aerial Images. In 2025 International Conference on Communication, Computer, and Information Technology (IC3IT) (pp. 1-6). IEEE.

47. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., et al. (2022). Retrieval-augmented generation for knowledge-intensive NLP tasks. Journal of Machine Learning Research, 23, 1–48.

48. Kandpal, N., Deng, H., Roberts, A., Wallace, E., & McCallum, A. (2023). Large language models struggle to learn long-tail knowledge. In Proceedings of the 40th International Conference on Machine Learning.

49. Kolla, S. H., & Mangala, N. (2025). DESIGNING AUTONOMOUS LLM AGENT FRAMEWORKS USING GEN AI PIPELINES TO ENHANCE CUSTOMER SERVICE MANAGEMENT AND KNOWLEDGE WORKFLOWS. Lex Localis-Journal of Local Self-Government, 23, 9719-9733.

50. He, Y., Bian, S., Chen, L., Hui, Y., Lentz, M., Li, B., Liu, F., et al. (2024). AI-native computing for large generative models: Opportunities and challenges. arXiv.

51. Golda, A., Chamola, V., Hassija, V., & Sikdar, B. (2024). Security and privacy challenges in generative artificial intelligence: Current trends and future directions. IEEE Access.

52. Nabende, P., & Wanyama, T. (2008). An expert system for diagnosing heavy-duty diesel engine faults. In Advances in computer and information sciences and engineering (pp. 384-389). Dordrecht: Springer Netherlands.

53. Deng, S., Dustdar, S., Yin, J., & Zomaya, A. Y. (2024). Service-oriented cloud-native computing: Architecture, orchestration, and future research directions. Proceedings of the IEEE.

54. Marin, E., Perino, D., & Di Pietro, R. (2022). Security challenges and research opportunities in serverless cloud-native systems. Journal of Cloud Computing.

55. Mangalampalli, Bindu Madhavi, Sasi Kumar Kolla, Velangani Divya Vardhan Kumar Bandi, Uday Surendra Yandamuri, and PR Sudha Rani. "Designing intelligent healthcare ecosystems through adaptive data integration and autonomous learning systems." Vascular and Endovascular Review 8, no. 20s (2025): 330-347.

56. Andreoni, M., Lunardi, W. T., Lawton, G., & Thakkar, S. (2024). Enhancing autonomous system security and resilience with generative AI: A comprehensive survey. IEEE Access, 12, 109470–109493.

57. Golda, A., Mekonen, K., Pandey, A., Singh, A., Hassija, V., Chamola, V., & Sikdar, B. (2024). Privacy and security concerns in generative AI: A comprehensive survey. IEEE Access, 12, 48126–48144.

58. Mangalampalli, B. M., & Kolla, S. K. (2025). Large Language Models for Automated Healthcare Data Dictionary Generation and Maintenance. Vascular and Endovascular Review, 8(20s), 363-375.

59. Deng, S., Zhao, H., Huang, B., Zhang, C., Chen, F., Deng, Y., Yin, J., Dustdar, S., & Zomaya, A. Y. (2024). Cloud-native computing: A survey from the perspective of services. Proceedings of the IEEE, 112(1), 12–46.

60. Syed, N. F., Shah, S. W., Shaghaghi, A., Anwar, A., Baig, Z., & Doss, R. (2022). Zero trust architecture (ZTA): A comprehensive survey. IEEE Access, 10, 57143–57179.

61. Ashokkumar, S., & Amistapuram, K. (2025, October). Attention-Guided Spatial Temporal Framework for Deepfake Detection on Social Video Platforms. In 2025 International Conference on Communication, Computer, and Information Technology (IC3IT) (pp. 1-6). IEEE.

62. Poirrier, A., Cailleux, L., & Clausen, T. H. (2025). Is trust misplaced? A zero-trust survey. Proceedings of the IEEE, 113(1), 5–39.

63. Lu, Y., Bian, S., Chen, L., He, Y., Hui, Y., Lentz, M., Li, B., Liu, F., Li, J., Liu, Q., Liu, R., Liu, X., Ma, L., Rong, K., Wang, J., Wu, Y., Wu, Y., Zhang, H., Zhang, M., Zhang, Q., Zhou, T., & Zhuo, D. (2024). Computing in the era of large generative models: From cloud-native to AI-native. arXiv.

64. Reddy, V. A. R., & Kolla, S. K. (1984). Infrastructure-As-Code Practices For Regulated Healthcare Cloud Environments. Metallurgical and Materials Engineering, 30 (4), 1028–1042.

65. OpenAI. (2023). GPT-4 technical report. arXiv:2303.08774.

66. Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., et al. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288.

67. Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D., et al. (2024). Mixtral of experts. arXiv:2401.04088.

68. Garapati, R. S., & Kanna, S. R. A Digital Twin‑Enabled Predictive Maintenance Framework Leveraging Multi‑Agent Reinforcement Learning and Industrial IoT Data.

69. Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., et al. (2023). A survey of large language models. arXiv.

70. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., et al. (2022). On the opportunities and risks of foundation models. arXiv.

71. Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., et al. (2022). Emergent abilities of large language models. Transactions on Machine Learning Research.

72. Davuluri, P. N. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems.

73. Chiang, W.-L., Zheng, L., Sheng, Y., Li, T., Zhuang, Z., Wu, Y., et al. (2023). Vicuna: An open-source chatbot impressing GPT-4 with 90% ChatGPT quality. arXiv.

74. Weinberg, A. I., & Cohen, K. (2024). Zero trust implementation in the emerging technologies era: Survey. arXiv.

75. Thijsman, J., Sebrechts, M., De Turck, F., & Volckaert, B. (2024). Trusting the cloud-native edge: Remotely attested Kubernetes workers. arXiv.

76. LEBCIR, I., Shah, C. A., & Appa Rao Nagubandi, D. S. M. D. FinTech and Financial Inclusion: Empirical Evidence from Emerging Markets.

77. Nahar, N., Andersson, K., Schelen, O., & Saguna, S. (2024). A survey on zero trust architecture: Applications and challenges of 6G networks. IEEE Access.

78. Bandi, V. D. V. K. AI-Based Anomaly Detection Frameworks in Distributed Enterprise Data Systems.

79. Patel, D., Raut, G., Cheetirala, S. N., Nadkarni, G. N., Freeman, R., Glicksberg, B. S., Klang, E., & Timsina, P. (2024). Cloud platforms for developing generative AI solutions: A scoping review of tools and services. arXiv.

80. Marin, E., Perino, D., & Di Pietro, R. (2022). Serverless computing: A security perspective. Journal of Cloud Computing, 11, Article 69.

81. GARAPATI, R. S. SYNERGETIC INTELLIGENCE Converging AI, Cloud, IoT, and Smart Automation for Real-Time Futures. CANEDA GLOBAL JOURNAL GROUP.

82. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al. (2022). Retrieval-augmented generation for knowledge-intensive NLP tasks. Journal of Machine Learning Research, 23, 1–48.

83. Kandpal, N., Deng, H., Roberts, A., Wallace, E., & McCallum, A. (2023). Large language models struggle to learn long-tail knowledge. In Proceedings of the 40th International Conference on Machine Learning.

84. Reddy, V. A. R. (2025). Journal of Rare Cardiovascular Diseases. Health, 5(3), 402-422.

Additional Files

Published

2025-03-14

Data Availability Statement

none

Most read articles by the same author(s)

Similar Articles

1-10 of 25

You may also start an advanced similarity search for this article.