Governance-Driven MLOps Using Databricks Medallion Architecture
DOI:
https://doi.org/10.5281/zenodo.21455258Keywords:
Data Governance Integration, MLOps Frameworks, Data Management Maturity, AI Data Governance, Machine Learning Operations, Medallion Architecture, DataForge Framework, Databricks MLOps, Data Quality Management, Regulatory Compliance in Data, Trusted Analytics Systems, Data Lifecycle Management, Model Governance, Data and Model Operations, Data Reliability Engineering, Governance Policy Layers, Scalable Data Operations, Data Risk Management, End-to-End MLOps, Enterprise Data Platforms.Abstract
Integrating Data Governance and MLOps support within data management is a recent requirement accelerating in the current decade. Several factors drive the demand for data governance: regulatory compliance, trust in analytics and machine learning insights, the need for data reuse controlling operational risks, and adequate data quality. The multiple dimensions of data management complexity require specialization, with increasingly higher levels of reliability. A maturity model derived from cybersecurity concepts divides data management in data management and data science operations with horizontal and policy layers spanning the intensity and breadth of data management, relying on appropriate degrees of specialization are.
MLOps methodologies are still maturing while the requirements from different domains push towards the definition of more specific solution, including most of the aspects demanded from modern-day data management operations. The integration of Data Governance and MLOps strategies for Machine Learning workloads is an evident necessity in the current panorama of data operations. The definition of the Bronze layer for the DataForge helps the coverage of Data Governance needs in Machine Learning operations, presenting a Medallion Architecture supported by Databricks tools able to provide end-to-end coverage of Data Governance Data Management processes and MLOps requirements from Data and Model development until Data and Models production.
References
1. Amou Najafabadi, F., Bogner, J., Gerostathopoulos, I., & Lago, P. (2024). An analysis of MLOps architectures: A systematic mapping study. Proceedings of the European Conference on Software Architecture, 112–129.
2. Sanepalli, U. R. (2024). Operationalizing MLOps with Databricks pipelines: Scalable machine learning in cloud environments. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 10(6), 2544–2552.
3. Kumar, S. S., Garapati, R. S., Segireddy, A. R., Kalisetty, S., Inala, R., & Nagabhyru, K. C. (2026). Hybrid Deep Neural Network–DevOps Pipeline Optimization for Risk Prediction in Cloud-Native Workers’ Compensation Platforms. In 2026 IEEE International Conference for Convergence in Computing Technology (I3CTCON) (pp. 1–6). IEEE. 2026 IEEE International Conference for Convergence in Computing Technology (I3CTCON). https://doi.org/10.1109/i3ctcon68242.2026.11507219
4. Brahmandam, L. M. K. (2024). An empirical evaluation of the medallion architecture on Databricks and Apache Spark with Snowflake. International Journal of AI Big Data Computational and Management Studies, 5(3), 197–206.
5. Sudha Rani, P. R., Amistapuram, K., Pamisetty, V., Singireddy, S., Kummari, D. N., & Sheelam, G. K. (2025). Hybrid Knowledge Graph–Deep Learning Framework for Automated Exception Handling and Investigation in Complex Insurance Claims. In 2025 IEEE 3rd Global Conference on Wireless Computing and Networking (GCWCN) (pp. 1–6). IEEE. 2025 IEEE 3rd Global Conference on Wireless Computing and Networking (GCWCN). https://doi.org/10.1109/gcwcn66157.2025.11448301
6. Bianchi, L., Romano, G., & Conti, F. (2024). Governed cloud AI for financial services: MLOps, foundation models, and trustworthy AI analytics. Journal of AI Analytics and Applications, 2(2), 55–71.
7. Salami, S. (2025). Hub star modeling 2.0 for medallion architecture. arXiv Preprint.
8. Lopes, C. L. V., Pitta, J. M., Belém, F., Alves, G., & Martins, F. V. C. (2026). Engineering AI agents for clinical workflows: A case study in architecture, MLOps, and governance. arXiv Preprint.
9. Kreuzberger, D., Kühl, N., & Hirschl, S. (2024). Machine learning operations: Architecture and lifecycle management. Journal of Systems and Software, 214, 111942.
10. Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., & Young, M. (2024). Hidden technical debt in machine learning systems revisited. Communications of the ACM, 67(2), 84–93.
11. Kumar, S. S., Garapati, R. S., Segireddy, A. R., Kalisetty, S., Inala, R., & Nagabhyru, K. C. (2026). Hybrid Deep Neural Network–DevOps Pipeline Optimization for Risk Prediction in Cloud-Native Workers’ Compensation Platforms. In 2026 IEEE International Conference for Convergence in Computing Technology (I3CTCON) (pp. 1–6). IEEE. 2026 IEEE International Conference for Convergence in Computing Technology (I3CTCON). https://doi.org/10.1109/i3ctcon68242.2026.11507219
12. Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2024). ML test score: A rubric for production readiness. IEEE Software, 41(1), 34–42.
13. Zaharia, M., Chen, A., Davidson, A., Ghodsi, A., Hong, S. A., Konwinski, A., & Xin, R. (2024). Lakehouse architecture for enterprise AI systems. VLDB Journal, 33(1), 45–68.
14. Armbrust, M., Huai, Y., Liang, C., Xin, R., & Zaharia, M. (2024). Delta Lake for reliable machine learning pipelines. Data Engineering Bulletin, 47(1), 22–37.
15. Polyzotis, N., Roy, S., Whang, S., & Zinkevich, M. (2024). Data validation for machine learning. Proceedings of VLDB, 17(4), 789–803.
16. Kumar, A., Boehm, M., & Yang, J. (2024). Efficient feature stores for machine learning operations. SIGMOD Record, 53(1), 26–39.
17. Krishnan, M., Aitha, A. R., Amistapuram, K., Nandan, B. P., Kaulwar, P. K., & Singireddy, J. (2025). Human-in-the-Loop Hybrid Neuro-Symbolic AI Model for Reliable Data Engineering in High-Stakes Industrial Systems. In 2025 IEEE 3rd Global Conference on Wireless Computing and Networking (GCWCN) (pp. 1–7). IEEE. 2025 IEEE 3rd Global Conference on Wireless Computing and Networking (GCWCN). https://doi.org/10.1109/gcwcn66157.2025.11448516
18. Nahar, K., & Hasan, M. (2024). Governance-aware feature engineering in enterprise MLOps. IEEE Access, 12, 55432–55448.
19. Ribeiro, M. T., Singh, S., & Guestrin, C. (2024). Explainability in regulated AI systems. AI Magazine, 45(2), 18–32.
20. Molnar, C. (2024). Interpretable machine learning for production systems. Journal of Artificial Intelligence Research, 80, 201–228.
21. Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., & Gebru, T. (2024). Model cards for model governance. Communications of the ACM, 67(4), 52–61.
22. Holland, S., Hosny, A., Newman, S., Joseph, J., & Chmielinski, K. (2024). Dataset documentation for responsible AI governance. Nature Machine Intelligence, 6(4), 312–321.
23. Gupta, D. K., Purushotham, K., Dheer, G., P, S., Gottimukkala, V. R. R., & Kapoor, S. (2025). Semantic Feature Learning Using Transformer-Based Deep Neural Networks. In 2025 IEEE 5th International Conference on ICT in Business Industry & Government (ICTBIG) (pp. 1–6). IEEE. 2025 IEEE 5th International Conference on ICT in Business Industry & Government (ICTBIG). https://doi.org/10.1109/ictbig68706.2025.11323734
24. NIST. (2024). Artificial intelligence risk management framework 1.1. National Institute of Standards and Technology.
25. ISO/IEC. (2024). ISO/IEC 42001: Artificial intelligence management systems. International Organization for Standardization.
26. European Commission. (2024). EU Artificial Intelligence Act. Publications Office.
27. Redman, T. C. (2024). Data quality for trusted AI systems. Harvard Data Science Review, 6(1), 1–15.
28. Bhavani, B. D., SR, S., Loganathan, R., & Nagaraj, S. (2026, April). Evolutionary Gravitational Neocognitron Neural Network, Snow Leopard Optimization and Deep Graph Reinforcement Learning for Routing Protocol in WSN. In 2026 2nd International Conference on Intelligent Systems and Computational Networks (ICISCN) (pp. 1-8). IEEE.
29. Rekatsinas, T., Chu, X., Ilyas, I., & Ré, C. (2024). HoloClean for enterprise data quality pipelines. VLDB Journal, 33(2), 321–338.
30. Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2024). Automating data validation in ML pipelines. Machine Learning Systems, 2(1), 1–21.
31. Bayram, O., & Korkmaz, M. (2024). Governance-aware CI/CD for machine learning deployment. Software Quality Journal, 32(2), 689–711.
32. Bandi, V. D. V. K. Autonomous Data Platforms: Converging AI, MLOps, and Cloud Engineering for Digital.
33. Bosch, J., & Olsson, H. (2024). Continuous deployment of AI systems. Journal of Software Evolution and Process, 36(3), e2521.
34. Rahman, A., & Gupta, S. (2024). Enterprise MLOps maturity assessment framework. IEEE Access, 12, 78220–78238.
35. Li, H., Xu, P., & Wang, J. (2024). Data lineage tracking in cloud data platforms. Future Generation Computer Systems, 154, 233–247.
36. Kolla, S. H., & Peddi, R. K. (2024). Designing Governance-Aligned GenAI Pipelines Using Small Language Models for Enterprise Workflow Intelligence. International Journal of Science, Research and Technology, 7(6), 13256-13268.
37. Verma, P., & Saini, R. (2024). Metadata governance in AI pipelines. Information Systems Frontiers, 26(4), 1301–1319.
38. Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2024). Concept drift in machine learning systems. ACM Computing Surveys, 57(2), 1–39.
39. Baier, L., Jöhren, F., & Seebacher, S. (2024). Challenges in drift monitoring for ML production. Expert Systems with Applications, 249, 123874.
40. Paleyes, A., Urma, R., & Lawrence, N. (2024). Challenges in deploying ML. Patterns, 5(2), 100932.
41. Miao, H., Li, A., Davis, L., & Deshpande, A. (2024). AI observability for enterprise platforms. IEEE Software, 41(3), 58–66.
42. Lakhanpal, S., Tunjungsari, H. K., Raj, S., Chukka, A., Dawadi, D., & Mangalampalli, B. M. (2026, April). The Digital Transformation of Healthcare: Balancing Opportunities, Risks, and Ethical Considerations. In 2026 International Conference on Multidisciplinary Innovations For Smart & Sustainable Future (MISSF) (pp. 1-6). IEEE.
43. Xin, R., Franklin, M., Stoica, I., & Zaharia, M. (2024). Unified analytics and AI on lakehouse systems. VLDB Endowment, 17(5), 1120–1133.
44. Kelleher, J. D., & Tierney, B. (2024). Governance in AI lifecycle management. AI and Ethics, 4(2), 411–425.
45. Batarseh, F., & Yang, R. (2024). Trustworthy MLOps in enterprise systems. Computers, 13(2), 44.
46. Singh, N., & Reddy, K. (2025). Secure model registry design for governed MLOps. Journal of Cloud Computing, 14(1), 19.
47. Raj, S., Prashanthi, P., Kolla, S. K., Bandaru, S., & Bhardwaj, N. (2026, April). Lightweight Cryptographic Schemes for Resource-Constrained IoT and Healthcare Devices. In 2026 International Conference on Multidisciplinary Innovations For Smart & Sustainable Future (MISSF) (pp. 1-6). IEEE.
48. Wang, X., Liu, Y., & Chen, H. (2025). Scalable feature stores in lakehouse ecosystems. IEEE Transactions on Big Data, 11(1), 88–103.
49. Patel, D., & Shah, R. (2025). Databricks Unity Catalog for enterprise governance. International Journal of Cloud Applications, 9(1), 34–49.
50. Inala, R., Garapati, R. S., Aitha, A. R., Komaragiri, V. B., Gottimukkala, V. R. R., Recharla, M., ... & Varri, D. B. S. (2026). U.S. Patent Application No. 19/389,108.
51. Fernandes, P., & Costa, M. (2025). Data contracts for AI pipelines. Software Practice and Experience, 55(4), 921–939.
52. Roy, S., & Das, P. (2025). Automated retraining orchestration in MLOps. Expert Systems, 42(1), e13602.
53. Kim, J., & Park, S. (2025). Responsible AI lifecycle orchestration. AI Ethics Review, 3(1), 55–73.
54. Mangala, S. B. G. N., Kolla, S. K., Segireddy, A. R., & Yandamuri, U. S. (2026). Designing Intelligent, Scalable Data Ecosystems for Precision Healthcare: A Cloud-Native Approach to Predictive and Automated Decision Systems. Advanced Engineering Sciences, 58(1), 2138-2152.
55. Gupta, V., & Mehta, S. (2025). Lakehouse security patterns for AI workloads. Journal of Information Security, 18(2), 145–162.
56. Zhao, T., & Huang, Y. (2025). Auditability in enterprise machine learning pipelines. Decision Support Systems, 185, 114298.
57. Davuluri, P. N. (2026). Autonomous Compliance Systems: AI, Event Streaming, and the Future of Financial Crime Prevention. Journal of Informatics Education and Research.
58. Ahmed, M., & Rahman, T. (2025). Model lineage frameworks for regulated AI. Knowledge-Based Systems, 310, 112458.
59. Chandra, R., & Banerjee, S. (2025). Policy-driven governance for AI systems. IEEE Access, 13, 22677–22695.
60. Brown, K., & Wilson, J. (2025). Automated compliance checks in MLOps. Software Quality Professional, 27(3), 14–28.
61. Mattaparthi, R. (2022). From Raw Sensor to Business Signal: An Azure Databricks Framework for Diesel Engine Performance Analytics Across Global Fleet Operations. Journal of Artificial Intelligence & Cloud Computing, 1(4), 1. https://doi.org/10.47363/jaicc/2022(1)529
62. Oliveira, L., & Sousa, R. (2025). AI governance metrics for model reliability. Applied Artificial Intelligence, 39(5), 410–431.
63. Park, D., & Lee, J. (2025). Delta Live Tables for governed streaming ML. Journal of Big Data Engineering, 7(2), 98–114.
64. Mangala, N. (2026). Responsible AI Data Architecture: Embedding GDPR and PII Compliance into MLOps Pipelines at Enterprise Scale. Canadian Journal of Marketing Research, 16(1), 107-124.
65. Narayanan, A., & Bhat, R. (2025). Feature lineage across medallion layers. Data & Knowledge Engineering, 159, 102411.
66. Chen, Y., & Luo, F. (2025). Enterprise lakehouse optimization for AI workloads. Future Internet, 17(4), 72.
67. Mangalampalli, B. M., & Kolla, T. (2026). FHIR-Based Interoperability Frameworks For Real-Time Healthcare Data Exchange: Architecture Patterns And Performance Optimization. International Journal Of Advances in Signal and Image Sciences, 1514-1536.
68. Singh, P., & Thomas, J. (2025). Governance-first model deployment using MLflow. Cloud Computing Advances, 8(2), 66–81.
69. Huang, Z., & Lin, C. (2025). Federated MLOps governance patterns. Journal of Network and Computer Applications, 235, 103901.
70. Krishnan, M., Bandi, V. D. V. K., Mangala, N., Kolla, S. H., & Mangalampalli, B. M. Engineering Intelligent Cloud-Native Data Ecosystems for Predictive Decision-Making in Industry.
71. Kumar, S., & Jain, P. (2025). Data observability in production AI systems. Information Sciences, 689, 120–138.
72. Das, R., & Iyer, S. (2026). Adaptive governance policies for autonomous ML systems. IEEE Transactions on AI, 7(1), 44–59.
73. Kulkarni, K. (2026). Building Context-Aware Enterprise Intelligence: A Secure and Scalable Architecture for AI-Augmented Data Systems. Minnesota Journal of Business Law and Entrepreneurship, (1), 1315-1336.
74. Romero, F., & Silva, D. (2026). Self-healing MLOps pipelines in enterprise AI. Artificial Intelligence Review, 59(2), 2101–2125.
75. Hassan, M., & Ali, N. (2026). Governance-aware AI agents in production environments. Journal of Intelligent Systems, 35(1), 90–109.
76. Nair, A., & George, R. (2026). AI audit trails using lakehouse metadata. Journal of Data Governance, 4(1), 1–17.
77. Sharma, R., & Kulkarni, P. (2026). Automated policy enforcement in Databricks pipelines. International Journal of Cloud Engineering, 11(1), 23–39.
78. Vega, M., & Torres, L. (2026). End-to-end explainability for governed MLOps. Expert Systems with Applications, 278, 127021.
79. Rani, S., Thavara, S. S., Kr, A., Naguband, A. R., Garg, A., & Vajpayee, A. Financialization of Sustainability: ESG Signalling, Market Valuation and the New Economics of Corporate Legitimacy.
80. Patel, K., & Desai, M. (2026). Continuous compliance monitoring for AI platforms. Computers & Security, 141, 103744.
81. Yang, Q., & Zhou, L. (2026). Metadata-centric MLOps governance. Information Systems, 129, 102459.
82. Rao, V., & Srinivasan, S. (2026). Databricks lakehouse governance patterns for enterprise AI. Journal of Enterprise Architecture, 22(1), 33–51.
83. Madhavi, B., Kolla, S. K., & Gari, V. A. K. R. R. (2026). Comment on" Using mobile applications for body composition analysis: A technical review of an artificial intelligence-based tool". Clinical nutrition ESPEN, 103365.
84. Martins, H., & Pereira, J. (2026). Policy-driven AI lifecycle orchestration. AI and Society, 41(2), 441–458.
85. Edwards, T., & Morgan, P. (2026). Governance-driven MLOps reference architecture for regulated industries. Journal of Machine Learning Systems, 5(1), 1–26.
Additional Files
Published
Issue
Section
License
Copyright (c) 2026 European Advanced Journal for Science & Engineering (EAJSE)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Articles published in the European Advanced Journal for Science & Engineering (EAJSE) are made freely available online immediately upon publication under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0). This license permits unrestricted use, distribution, and reproduction in any medium or format, provided the original work is properly cited. Authors retain copyright of their work. By submitting to EAJSE, authors grant the journal the right of first publication. For details, visit: https://creativecommons.org/licenses/by/4.0/