A Distributed Execution Framework for Enterprise Decision Systems

Authors

  • Katarzyna Nowak Author

Keywords:

Decision intelligence, enterprise governance, distributed AI, execution frameworks, orchestration, policy governance, scalability ,Scalable Decision Intelligence,Enterprise Governance,Distributed Artificial Intelligence,AI Execution Frameworks,Intelligent Decision Support Systems,Distributed Machine Learning,Enterprise AI Architecture,Governance Automation,Adaptive Decision Analytics,Explainable and Trustworthy AI.

Abstract

Enterprise governance of Decision Intelligence enables organizations to make decisions at scale, responsibly, in a world full of both challenges and opportunities. The adoption of Distributed AI Execution Frameworks has made such scalable decision intelligence achievable in practice; however, governing these distributed frameworks properly remains an open problem. Governance mechanisms are essential for mitigating the associated risks and ensuring that accountability and transparency requirements are met. Policy-driven decision governance addresses the specification, enforcement, compliance, override, and audit of decision processes — enabling enterprises to realize the full potential of this approach.

Decision Intelligence (DI) reflects a growing effort among organizations to address the many factors that shape decision-making in a holistic way. As a metadiscipline, DI brings together Data-Driven Decision-Making and Machine Learning-Based Decision-Making, ensuring that neither happens in isolation but instead forms part of a coherent, well-considered strategy. Enterprise governance of DI thus enables decision-making that is scalable, transparent, auditable, and responsible — a capability sorely needed in a world that demands swift, competent decisions at every level of society.

References

1. Li, T., Sahu, A. K., Talwalkar, A., & Smith, V. (2020). Federated learning: Challenges, methods, and future directions. IEEE Signal Processing Magazine, 37(3), 50–60.

2. Rajbhandari, S., Rasley, J., Ruwase, O., & He, Y. (2020). ZeRO: Memory optimizations toward training trillion parameter models. Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, 1–16.

3. Lepikhin, D., Lee, H., Xu, Y., Chen, D., Fırat, O., Huang, Y., Krikun, M., Shazeer, N., & Chen, Z. (2021). GShard: Scaling giant models with conditional computation and automatic sharding. International Conference on Learning Representations.

4. Garapati, R. S., Aitha, A. R., Yandamuri, U. S., Gottimukkala, V. R. R., Nagubandi, A. R., & Kolla, S. H. (2026, March). Cloud-Native Orchestration of Multi-Counterparty Derivatives and Collateral in Manufacturing Enterprises via AI-Assisted Financial Audit Engines. In 2026 IEEE International Conference on AI Engineering and Innovations (AIEI) (pp. 1-6). IEEE.

5. Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120), 1–39.

6. Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., Gonzalez, J. E., & Stoica, I. (2022). Alpa: Automating inter- and intra-operator parallelism for distributed deep learning. Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation, 559–578.

7. Lai, F., Dai, Y., Singapuram, S. S., Liu, J., Zhu, X., Madhyastha, H. V., & Chowdhury, M. (2022). FedScale: Benchmarking model and system performance of federated learning at scale. Proceedings of the 39th International Conference on Machine Learning, 162, 11814–11827.

8. Elkordy, A. R., Ezzeldin, Y. H., Han, S., Sharma, S., He, C., Mehrotra, S., & Avestimehr, S. (2023). Federated analytics: A survey. APSIPA Transactions on Signal and Information Processing, 12(1).

9. Segireddy, A. R., Davuluri, P. N., Mangala, N., & Nagubandi, A. R. (2026, April). Self-Healing Cloud-Native Quorum Systems: AI-Enabled DevOps for Zero-Downtime Financial Payments. In 2026 International Conference on NextGen Data Science and Analytics (ICNDSA) (pp. 1-9). IEEE.

10. Dean, P., & Porter, B. (2021). The design space of emergent scheduling for distributed execution frameworks. Proceedings of the Symposium on Software Engineering for Adaptive and Self-Managing Systems, 186–195.

11. Thamsen, L. (2020). Mary, Hugo, and Hugo*: Learning to schedule distributed data-parallel processing jobs on shared clusters. Concurrency and Computation: Practice and Experience, 32(20), e5823.

12. Bergui, M., Najah, S., & Nikolov, N. S. (2021). A survey on bandwidth-aware geo-distributed frameworks for big-data analytics. Journal of Big Data, 8, Article 40.

13. Li, J., Wang, Z., & Zhang, X. (2021). Octopus-DF: Unified DataFrame-based cross-platform data analytic system. Parallel Computing, 105, 102879.

14. Widanage, C., et al. (2024). Supercharging distributed computing environments for high-performance data engineering. Frontiers in High Performance Computing, 2.

15. Isukapalli, S., & Srirama, S. N. (2024). A systematic survey on fault-tolerant solutions for distributed data analytics: Taxonomy, comparison, and future directions. Computer Science Review, 53, 100660.

16. Zhang, C., Li, Y., & Li, X. (2020). Distributed deep learning with communication-efficient optimization. IEEE Transactions on Parallel and Distributed Systems, 31(10), 2384–2397.

17. Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., D’Oliveira, R. G. L., Eichner, H., El Rouayheb, S., Evans, D., Gardner, M. J., Garrett, Z., Gascón, A., Ghazi, B., Gibbons, P. B., ... Zhao, S. (2021). Advances and open problems in federated learning. Foundations and Trends in Machine Learning, 14(1–2), 1–210.

18. Aitha, A. R. (2026). Explainable Agentic AI Framework for Automated Insurance Fraud Detection and Predictive Risk Intelligence. International Journal of Advanced Research in Computer Science & Technology (IJARCST), 9(3), 877-888.

19. Li, Q., Wen, Z., Wu, Z., Hu, S., Wang, N., He, B., & Li, J. (2020). A survey on federated learning systems: Vision, hype and reality for data privacy and protection. IEEE Transactions on Knowledge and Data Engineering.

20. Bonawitz, K., Kairouz, P., McMahan, H. B., Rambhatla, S., et al. (2022). Towards federated learning at scale: System design. Proceedings of Machine Learning and Systems, 4, 374–388.

21. Verbraeken, J., Wolting, M., Katzy, J., Kloppenburg, J., Verbelen, T., & Rellermeyer, J. S. (2020). A survey on distributed machine learning. ACM Computing Surveys, 53(2), Article 30.

22. Shi, W., Cao, J., Zhang, Q., Li, Y., & Xu, L. (2020). Edge computing: Vision and challenges. IEEE Internet of Things Journal, 7(4), 3164–3178.

23. Satyanarayanan, M. (2020). The emergence of edge computing. Computer, 53(1), 30–39.

24. Inala, R., Garapati, R. S., Aitha, A. R., Komaragiri, V. B., Gottimukkala, V. R. R., Recharla, M., ... & Varri, D. B. S. (2026). U.S. Patent Application No. 19/389,108.

25. Zhou, Z., Chen, X., Li, E., Zeng, L., Luo, K., & Zhang, J. (2020). Edge intelligence: Paving the last mile of artificial intelligence with edge computing. Proceedings of the IEEE, 107(8), 1738–1762.

26. Zhou, X., Chen, H., & Wang, X. (2021). Distributed machine learning systems: A survey of architectures, algorithms, and applications. IEEE Access, 9, 109897–109916.

27. Ben-Nun, T., & Hoefler, T. (2020). Demystifying parallel and distributed deep learning: An in-depth concurrency analysis. ACM Computing Surveys, 52(4), Article 65.

28. Sergeev, A., & Del Balso, M. (2020). Horovod: Fast and easy distributed deep learning in TensorFlow. Journal of Machine Learning Research, 21(79), 1–10.

29. Peng, Y., Zhu, Y., Bao, Y., Chen, Y., Wu, C., Guo, C., & Lin, Q. (2020). A generic communication-efficient parallelization method for distributed deep learning. Proceedings of the 2020 USENIX Annual Technical Conference, 1–14.

30. Huang, C.-C., Jin, G., & Dai, J. (2020). SwapAdvisor: Pushing deep learning beyond the GPU memory limit. Proceedings of the 27th International Conference on Architectural Support for Programming Languages and Operating Systems, 367–380.

31. Zhuang, Y., Zheng, L., Zhao, Z., Li, Z., & Stoica, I. (2023). On optimizing the communication of model parallelism. Proceedings of Machine Learning and Systems, 5, 1–14.

32. Yandamuri, U. S. (2026). Scalable Cloud-Based Intelligent Decision Systems Leveraging AI and Big Data for Industry-Specific Optimization. Minnesota Journal of Business Law and Entrepreneurship, (1), 584-601.

33. Li, Z., Zheng, L., Zhong, Y., Liu, V., Sheng, Y., Jin, X., Huang, Y., Chen, Z., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). AlpaServe: Statistical multiplexing with model parallelism for deep learning serving. Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation, 663–679.

34. Widanage, C., Perera, D., et al. (2024). Supercharging distributed computing environments for high-performance data engineering. Frontiers in High Performance Computing, 2.

35. Isukapalli, S., & Srirama, S. N. (2024). A systematic survey on fault-tolerant solutions for distributed data analytics: Taxonomy, comparison, and future directions. Computer Science Review, 53, 100660.

Additional Files

Published

2026-08-17

Data Availability Statement

None

How to Cite

A Distributed Execution Framework for Enterprise Decision Systems. (2026). European Data Science Journal (EDSJ), 4(03). https://esa-research.org/index.php/EDSJ/article/view/212

Most read articles by the same author(s)

Similar Articles

1-10 of 52

You may also start an advanced similarity search for this article.