Self-Governing Agents for Data Center Service Delivery

Authors

  • Katarzyna Nowak Author

Keywords:

Autonomy and control, enterprise data centers, governance, multi-tenancy, service delivery, trust .

Abstract

Agentic AI consists of autonomous systems capable of perceiving their environment, reasoning about it, taking actions to affect it, and learning from experience. Enterprise-grade data centers provide services to multiple users through multimodal workloads running on shared infrastructures. Data center ecosystems are multitenant environments with strict service-level agreements and compliance requirements. Agentic AI introduces autonomy at level 2 for service delivery optimization. Control loops close the consequent decision-making boundaries, and policies govern permission models and align with enterprise-defined boundaries. Existing solutions focus primarily on detecting and responding to anomalies rather than on full-automation solutions with a formal agentic definition whose control loops span the service delivery stack.

An enterprise-grade data-center ecosystem encompasses many telemetries that characterize the state of its components and the health of its service deliveries. Monitoring and observability enable the detection of deviations from expected behaviors. Such deviations necessitate actions to mitigate performance impact or security threats and, therefore, control loops to optimize service deliveries are essential. ML-based detection and prediction models trigger actions when replacements, scalings, or routing adjustments are needed. Fault-tolerance mechanisms (e.g., replication, checkpointing) are in place to allow system degradation during resource pull-outs, and agents monitor health to assure these processes remain invisible to users.

References

1. Asghari, A., Sohrabi, M. K., & Yaghmaee, F. (2020). A cloud resource management framework for multiple online scientific workflows using cooperative reinforcement learning agents. Computer Networks, 179, 107340.

2. Caviglione, L., Gaggero, M., Paolucci, M., & Ronco, R. (2021). Deep reinforcement learning for multi-objective placement of virtual machines in cloud datacenters. Soft Computing, 25, 12569–12588.

3. Duan, Y., Wan, J., Zhou, J., Cong, G., Rasheed, Z., & Hua, T. (2020). Reinforcement learning for rack-level cooling. In Mobile Wireless Middleware, Operating Systems and Applications (pp. 158–168). Springer.

4. Inala, R. Designing Scalable Technology Architectures for Customer Data in Group Insurance and Investment Platforms.

5. Hu, X., & Sun, Y. (2020). A deep reinforcement learning-based power resource management for fuel cell powered data centers. Electronics, 9(12), 2054.

6. Li, Y., Wen, Y., Tao, D., & Guan, K. (2020). Transforming cooling optimization for green data center via deep reinforcement learning. IEEE Transactions on Cybernetics, 50(5), 2002–2013.

7. Le, D. V., Wang, R., Liu, Y., Tan, R., Wong, Y. W., & Wen, Y. (2021). Deep reinforcement learning for tropical air free-cooled data center control. ACM Transactions on Sensor Networks, 17(3), 24:1–24:28.

8. Thein, T., Myo, M. M., Parvin, S., & Gawanmeh, A. (2020). Reinforcement learning based methodology for energy-efficient resource allocation in cloud data centers. Journal of King Saud University—Computer and Information Sciences, 32(10), 1127–1139.

9. Shaw, R., Howley, E., & Barrett, E. (2022). Applying reinforcement learning towards automating energy efficient virtual machine consolidation in cloud data centers. Information Systems, 107, 101722.

10. Hua, T., Wan, J., Jaffry, S., Rasheed, Z., Li, L., & Ma, Z. (2021). Comparison of deep reinforcement learning algorithms in data center cooling management: A case study. In 2021 IEEE International Conference on Systems, Man, and Cybernetics (pp. 3360–3365). IEEE.

11. Heimerson, A., Sjölund, J., Brännvall, R., Gustafsson, J., & Eker, J. (2022). Adaptive control of data center cooling using deep reinforcement learning. In 2022 IEEE International Conference on Autonomic Computing and Self-Organizing Systems Companion (pp. 57–62). IEEE.

12. Li, R., Cheng, Z., Lee, P. P. C., Wang, P., Qiang, Y., Lan, L., He, C., Lu, J., Wang, M., & Ding, X. (2021). Automated intelligent healing in cloud-scale data centers. In Proceedings of the 40th International Symposium on Reliable Distributed Systems (pp. 244–253). IEEE.

13. Xiao, X., Sun, J., & Yang, J. (2021). Operation and maintenance (O&M) for data center: An intelligent anomaly detection approach. Computer Communications, 178, 141–152.

14. Alonso, J., Orue-Echevarria, L., Osaba, E., López Lobo, J., Martinez, I., Diaz de Arcaya, J., & Etxaniz, I. (2021). Optimization and prediction techniques for self-healing and self-learning applications in a trustworthy cloud continuum. Information, 12(8), 308.

15. Belgacem, A., Mahmoudi, S., & Kihl, M. (2022). Intelligent multi-agent reinforcement learning model for resources allocation in cloud computing. Journal of King Saud University—Computer and Information Sciences, 34(6), 2391–2404.

16. Wang, S., Qin, L., Ma, C., & Wu, W. (2023). Research on overall energy consumption optimization method for data center based on deep reinforcement learning. Journal of Intelligent & Fuzzy Systems, 44(5), 7333–7349.

17. Davuluri, P. S. L. (2023). AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems (December 15, 2023).

18. Heimerson, A., Sjölund, J., Brännvall, R., Gustafsson, J., & Eker, J. (2022). Adaptive control of data center cooling using deep reinforcement learning. In Proceedings of the IEEE International Conference on Autonomic Computing and Self-Organizing Systems Companion. IEEE.

19. Tan, R., Wang, R., Le, D. V., Liu, Y., Wong, Y. W., & Wen, Y. (2023). Phyllis: Physics-informed lifelong reinforcement learning for data center cooling control. In Proceedings of the 14th ACM International Conference on Future Energy Systems (pp. 114–126). Association for Computing Machinery.

20. Raj, A., Perarnau, S., & Gokhale, A. (2023). A reinforcement learning approach for performance-aware reduction in power consumption of data center compute nodes. In Proceedings of the 2023 IEEE International Conference on Autonomic Computing and Self-Organizing Systems. IEEE.

21. Dong, D. (2023). Agent-based cloud simulation model for resource management. Journal of Cloud Computing, 12, Article 145.

22. Poltronieri, F., Stefanelli, C., Tortonesi, M., & Zaccarini, M. (2023). Reinforcement learning vs. computational intelligence: Comparing service management approaches for the cloud continuum. Future Internet, 15(11), 359.

23. Naug, A., Guillen, A., Luna Gutiérrez, R., Gundecha, V., Markovikj, D., Kashyap, L. D., Krause, L., Ghorbanpour, S., Mousavi, S., Babu, A. R., & Sarkar, S. (2023). PyDCM: Custom data center models with reinforcement learning for sustainability. In Proceedings of the 2023 IEEE International Conference on Autonomic Computing and Self-Organizing Systems. IEEE.

24. Guo, W., Tian, W., Ye, Y., Xu, L., & others. (2021). Cloud resource scheduling with deep reinforcement learning and imitation learning. IEEE Internet of Things Journal, 8(19), 14970–14983.

25. Qin, Y., Wang, H., Yi, S., & others. (2020). Virtual machine placement based on multi-objective reinforcement learning. Applied Intelligence, 50(8), 2370–2383.

26. Karthik, A., & others. (2021). Resource scheduling approach in cloud Testing as a Service using deep reinforcement learning algorithms. CAAI Transactions on Intelligence Technology, 6(4), 483–495.

27. Li, B., Wang, T., Yang, P., Chen, M., Yu, S., & Hamdi, M. (2022). Machine learning empowered intelligent data center networking: A survey. IEEE Communications Surveys & Tutorials, 24(4), 2655–2690.

28. Gill, S. S., Xu, M., Ottaviani, C., Patros, P., Bahsoon, R., Shaghaghi, A., Golec, M., Stankovski, V., Wu, H., Abraham, A., Singh, M., Mehta, H., Ghosh, S. K., Baker, T., Parlikad, A. K., Lutfiyya, H., Kanhere, S. S., Sakellariou, R., Dustdar, S., Rana, O., Brandic, I., & Uhlig, S. (2022). AI for next generation computing: Emerging trends and future directions. Internet of Things, 19, 100514.

29. Amistapuram, K. (2023). Privacy-Preserving Machine Learning Models for Sensitive Customer Data in Insurance Systems. Educational Administration: Theory and Practice, 29(4), 5950-5958.

30. Feng, Y., & Liu, F. (2023). Resource management in cloud computing using deep reinforcement learning: A survey. In Proceedings of the 10th Chinese Society of Aeronautics and Astronautics Youth Forum (pp. 635–643). Springer.

31. Sheng, J., Wang, L., Yang, F., Qiao, B., Dong, H., Wang, X., Jin, B., Wang, J., Qin, S., Rajmohan, S., Lin, Q., & Zhang, D. (2022). Learning cooperative oversubscription for cloud by chance-constrained multi-agent reinforcement learning. arXiv.

32. Yandamuri, U. S. (2023). An Intelligent Analytics Framework Combining Big Data and Machine Learning for Business Forecasting. Zenodo.

33. Huang, X., Banerjee, A., Chen, C.-C., Huang, C., Chuang, T. Y., Srivastava, A., & Cheveresan, R. (2021). Challenges and solutions to build a data pipeline to identify anomalies in enterprise system performance. arXiv.

34. Yao, C. (2021). Highly efficient memory failure prediction using Mcelog-based data mining and machine learning. arXiv.

Additional Files

Published

2023-12-05

Data Availability Statement

None

How to Cite

Self-Governing Agents for Data Center Service Delivery. (2023). European Data Science Journal (EDSJ), 1(01). https://esa-research.org/index.php/EDSJ/article/view/108

Most read articles by the same author(s)

Similar Articles

1-10 of 47

You may also start an advanced similarity search for this article.