Medallion-Based Feature Engineering for Real-Time MLOps
Keywords:
Feature stores, DataOps, Continuous Integration, Continuous Delivery, Cloud-Native platforms, Medallion Architecture, Climate Simulation, City of New York, NYC Citywide Climate Networks, NYC City Sensor Data, Medium Data, Continuous Data EngineeringAbstract
As large-scale machine learning in production becomes ubiquitous across various industries, both engineering and research-oriented teams are pushing the limits for faster and continuously retrained models. Nevertheless, building and running a model, even using a well-established MLOps framework, is often only a small part of the overall task. Feature engineering and feature selection typically consume the majority of both the data and ML time of a project. A Medallion Architecture specifically designed to assist in real-time feature engineering, along with the establishment of an enterprise Feature Store capability, can help to speed this process, integrating the latest insights from the data into models and putting both new feature and model within continuous integration pipelines.
The architecture incorporates key DataOps and MLOps best practices for target labeling, feature quality, and CI/CD pipelines, while enabling features to be built in parallel either as part of a dedicated MLOps modeling step or part of the standard product development process. ML pipelines and Feature Stores for Scalable Real-Time Feature Engineering cover Medallion Architecture and a large-scale Medallion Implementation, integrating best practices for Real-Time Streaming Data Ingestion and Processing as well as Continuous Integration and Continuous Delivery for Features. As a result, a Medallion Architecture designed for data scale considers the needs of engineering-focused feature engineering while incorporating horizontal scaling through sharding strategies aligned with both features to be built and cost objectives around exploratory data and feature selection efforts.
References
1. Akiba, T., Sano, S., Yanase, T., Ohta, T., & Koyama, M. (2022). Optuna: A next-generation hyperparameter optimization framework. ACM Transactions on Knowledge Discovery from Data, 16(1), 1–30.
2. Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., Nagappan, N., Nushi, B., & Zimmermann, T. (2021). Software engineering for machine learning: A case study. Communications of the ACM, 64(7), 56–65.
3. Mattaparthi, R. (2024). Transformer-Based Fault Diagnosis for Large-Scale Standby Power Generators: Partial Discharge Pattern Recognition at Hyperscale Data Center Installations. Journal of Computational Analysis and Applications (JoCAAA), 33(08), 8781-8799.
4. Armbrust, M., Ghodsi, A., Xin, R., Zaharia, M., Torres, J., & Das, T. (2021). Lakehouse: A new generation of open platforms that unify data warehousing and advanced analytics. CIDR Proceedings.
5. Biewald, L. (2021). Experiment tracking with Weights & Biases. Journal of Machine Learning Systems, 3(2), 45–57.
6. Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2021). The ML test score: A rubric for ML production readiness and technical debt reduction. IEEE Software, 38(5), 82–89.
7. Oladosu, S. A., Ige, A. B., Ike, C. C., Adepoju, P. A., Amoo, O. O., & Afolabi, A. I. (2022). Revolutionizing data center security: Conceptualizing a unified security framework for hybrid and multi-cloud data centers. Open Access Research Journal of Science and Technology, 5(2), 086-076.
8. Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2021). End-to-end object detection with transformers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12), 9361–9375.
9. Chen, T., Li, E., Zhang, H., & Zhao, Y. (2023). Streaming feature engineering for scalable real-time machine learning systems. IEEE Access, 11, 112456–112470.
10. Chowdhury, A., & Ward, M. (2022). Feature stores for machine learning: A survey of architectures and operational practices. Journal of Big Data, 9(1), 87.
11. Reddy, V. A. R. (2023). Predictive Healthcare Administration Using Advanced Payer Analytics and Population Health Data Engineering. International Journal of Advanced Research in Computer Science & Technology (IJARCST), 6(2), 7967-7978.
12. Cruz, L., & Rodrigues, M. (2023). Data quality management for feature engineering pipelines in cloud-native environments. Future Generation Computer Systems, 141, 180–194.
13. Kreuzberger, D., Kühl, N., & Hirschl, S. (2022). Machine learning operations (MLOps): Overview, definition, and architecture. ACM Computing Surveys, 55(13), 1–39.
14. Kolla, T. (2024). Intelligent Discovery and Governance of Healthcare Data Assets Through AI-Powered Catalog Architectures. International Journal of Emerging Trends in Engineering and Management Research, 9(4), 16083.
15. Kumar, A., Lee, J., & Patel, R. (2024). Real-time feature stores for scalable machine learning inference. Journal of Systems and Software, 209, 111892.
16. Li, Y., Zhang, H., & Wang, X. (2023). Automated feature engineering for streaming analytics using Apache Spark. Future Internet, 15(9), 298.
17. Paleyes, A., Urma, R., & Lawrence, N. D. (2022). Challenges in deploying machine learning: A survey of case studies. ACM Computing Surveys, 55(6), 1–29.
18. Inala, R. AI-Powered Investment Decision Support Systems: Building Smart Data Products with Embedded Governance Controls.
19. Polyzotis, N., Roy, S., Whang, S. E., & Zinkevich, M. (2021). Data lifecycle challenges in production machine learning. ACM SIGMOD Record, 50(3), 17–28.
20. Sato, Y., Kimura, T., & Nakamura, H. (2024). Lakehouse-based feature engineering pipelines for real-time AI applications. IEEE Access, 12, 45672–45689.
21. Schelter, S., Biessmann, F., Januschowski, T., Salinas, D., Seufert, S., Szarvas, G., & Taptunov, A. (2021). On challenges in machine learning model management. IEEE Data Engineering Bulletin, 44(1), 5–15.
22. Kolla, S. K. (2022). Engineering Healthcare Data Infrastructures for Predictive Clinical Analytics and Evidence-Based Decision Making. International Journal of Engineering & Extended Technologies Research (IJEETR), 4(5), 5370-5380.
23. Shankar, S., Garcia, P., & Williams, T. (2023). Continuous feature validation for MLOps pipelines. Information Systems, 117, 102185.
24. Singh, R., Gupta, A., & Verma, P. (2024). Medallion architecture for scalable feature engineering in cloud lakehouses. Journal of Cloud Computing, 13(1), 66.
25. Zaharia, M., Chen, A., Davidson, A., Ghodsi, A., Hong, S., Konwinski, A., Murching, S., Nykodym, T., Ogilvie, P., Parkhe, M., Xie, F., & Zumar, C. (2021). Data lakehouse: The definitive guide. O'Reilly Media.
26. Davuluri, P. N. (2019). Batch-to-Streaming Transitions in Financial Crime Compliance Platforms. International Journal Of Engineering And Computer Science, 8(12).
27. Zhou, Q., Wang, L., & Chen, X. (2024). Real-time MLOps with feature stores and streaming data pipelines. IEEE Access, 12, 87421–87438.
28. Abbas, A., Khan, S. U., Ahmed, M., & Khan, M. A. (2022). Machine learning operations (MLOps): State of the art, challenges, and future research directions. Journal of Systems Architecture, 127, 102540.
29. Mangalampalli, B. M. (2024). Transparent Intelligence Explainability Frameworks for AI-Driven Clinical Decision Support in Healthcare Business Intelligence. International Journal of Research Publications in Engineering, Technology and Management (IJRPETM), 7(3), 10566-10579.
30. Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., … Zheng, X. (2021). TensorFlow: Large-scale machine learning on heterogeneous systems. Software: Practice and Experience, 51(1), 1–25.
31. Bandi, V. D. V. K. (2024). Intelligent Data Platforms For Personalized Retail Analytics At Scale. Metallurgical and Materials Engineering, 30(4), 1011-1027.
32. Brundage, M., Avin, S., Clark, J., Toner, H., Eckersley, P., Garfinkel, B., Dafoe, A., Scharre, P., Zeitzoff, T., Filar, B., Anderson, H., Roff, H., Allen, G. C., Steinhardt, J., Flynn, C., Héigeartaigh, S. Ó., Beard, S., Belfield, H., Farquhar, S., & Amodei, D. (2022). Toward trustworthy AI development and deployment. AI Magazine, 43(2), 147–161.
33. Chen, J., Li, X., Wang, H., & Zhao, L. (2023). Cloud-native data engineering for scalable machine learning pipelines. Journal of Cloud Computing, 12(1), 84.
34. Dunning, T., & Friedman, E. (2021). Streaming architecture for large-scale feature engineering. IEEE Internet Computing, 25(5), 66–74.
35. Elbaz, D., Roy, A., & Kumar, S. (2024). Continuous feature monitoring in production machine learning systems. IEEE Access, 12, 55487–55502.
36. Erickson, N., Mueller, J., Shirkov, A., Zhang, H., Larroy, P., Li, M., & Smola, A. (2021). AutoGluon-Tabular: Robust and accurate AutoML for structured data. Expert Systems with Applications, 182, 115290.
37. Peddi, R. K. (2024). AI-Based Workforce Analytics for SLA Governance and Uptime Assurance in Data Centers. Journal of Computational Analysis and Applications (JoCAAA), 33(08), 8589-8601.
38. Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2021). A survey on concept drift adaptation. ACM Computing Surveys, 54(6), 1–37.
39. Huyen, C. (2022). Designing machine learning systems: An iterative process for production-ready applications. O'Reilly Media.
40. Isah, H., Abughofa, T., Mahfuz, S., Ajerla, D., Zulkernine, F., & Khan, S. (2022). A survey of distributed data stream processing frameworks. IEEE Access, 10, 31130–31156.
41. Jin, H., Lee, S., & Park, J. (2023). Scalable online feature engineering using Apache Spark Structured Streaming. Future Generation Computer Systems, 143, 204–218.
42. Kreps, J., Narkhede, N., & Rao, J. (2021). Kafka: A distributed messaging system for modern data pipelines. IEEE Software, 38(4), 58–66.
43. Yandamuri, U. S. (2024). AI-Driven Decision Support Systems for Operational Optimization in Hospitality Technology. Metallurgical and Materials Engineering.
44.
45. Kumar, P., Sharma, V., & Singh, R. (2023). Feature lineage tracking for enterprise MLOps. Journal of Big Data, 10(1), 94.
46. Li, H., Chen, X., & Zhou, Y. (2024). Data quality assessment for machine learning feature pipelines in lakehouse environments. Information Systems, 122, 102347.
47. Liu, Z., Wang, Y., & Xu, F. (2023). Metadata-driven feature engineering for cloud-native machine learning platforms. Future Internet, 15(11), 371.
48. Miao, Y., Li, P., & Zhao, H. (2022). Automated data validation for reliable machine learning pipelines. IEEE Access, 10, 94586–94599.
49. Nguyen, T., Tran, H., & Pham, D. (2024). Lakehouse-based architectures for real-time artificial intelligence applications. Journal of Cloud Computing, 13(1), 72.
50. Pamisetty, V., & Amistapuram, K. Smart Decision Support Systems For Dynamic Tax Policy Optimization Using Reinforcement Learning.
51. O'Leary, D. E. (2022). MLOps systems and governance: A review. Intelligent Systems in Accounting, Finance and Management, 29(3), 157–170.
52. Patel, K., Shah, N., & Desai, R. (2023). Feature engineering automation for streaming analytics using Apache Flink. Journal of Parallel and Distributed Computing, 176, 110–123.
53. Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2022). Automating large-scale data quality verification in machine learning pipelines. Proceedings of the VLDB Endowment, 15(12), 3470–3483.
54. Mangala, N. (2024). Leveraging Microsoft Fabric lakehouse as an AI-ready data platform for enterprise analytics. Journal of Information Systems Engineering and Management.
55. Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2022). Automating large-scale data quality verification in machine learning pipelines. Proceedings of the VLDB Endowment, 15(12), 3470–3483.
56. Amou Najafabadi, F., Bogner, J., Gerostathopoulos, I., & Lago, P. (2024). An analysis of MLOps architectures: A systematic mapping study. In Software Architecture: 18th European Conference (ECSA 2024) (Lecture Notes in Computer Science, Vol. 14889, pp. 69–85). Springer.
57. Kreuzberger, D., Kühl, N., & Hirschl, S. (2022). Machine learning operations (MLOps): Overview, definition, and architecture. ACM Computing Surveys, 55(13), 1–35.
58. Davuluri, P. N. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems.
59. Huyen, C. (2022). Designing Machine Learning Systems: An Iterative Process for Production-Ready Applications. O'Reilly Media.
60. Li, A., Ranganathan, B., Pan, F., Zhang, M., Xu, Q., Li, R., Raman, S., Shah, S. P., & Tang, V. (2023). Managed geo-distributed feature store: Architecture and system design. arXiv.
61. Zaharia, M., Chen, A., Davidson, A., Ghodsi, A., Hong, S., Konwinski, A., Murching, S., Nykodym, T., Ogilvie, P., Parkhe, M., Xie, F., & Zumar, C. (2024). Data lakehouse: The definitive guide. O'Reilly Media.
62. Armbrust, M., Ghodsi, A., Xin, R., Zaharia, M., Torres, J., & Das, T. (2021). Lakehouse: A new generation of open platforms that unify data warehousing and advanced analytics. Proceedings of CIDR.
63. Polyzotis, N., Roy, S., Whang, S. E., & Zinkevich, M. (2021). Data lifecycle challenges in production machine learning. ACM SIGMOD Record, 50(3), 17–28.
64. Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2021). The ML test score: A rubric for ML production readiness and technical debt reduction. IEEE Software, 38(5), 82–89.
65. Paleyes, A., Urma, R., & Lawrence, N. D. (2022). Challenges in deploying machine learning: A survey of case studies. ACM Computing Surveys, 55(6), 1–29.
66. Isah, H., Abughofa, T., Mahfuz, S., Ajerla, D., Zulkernine, F., & Khan, S. (2022). A survey of distributed data stream processing frameworks. IEEE Access, 10, 31130–31156.
67. Erickson, N., Mueller, J., Shirkov, A., Zhang, H., Larroy, P., Li, M., & Smola, A. (2021). AutoGluon-Tabular: Robust and accurate AutoML for structured data. Expert Systems with Applications, 182, 115290.
68. Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2021). A survey on concept drift adaptation. ACM Computing Surveys, 54(6), 1–37.
69. Dunning, T., & Friedman, E. (2021). Streaming Architecture: New Designs Using Apache Kafka and MapR Streams. O'Reilly Media.
70. Kreps, J., Narkhede, N., & Rao, J. (2021). Kafka: A distributed messaging system for modern data pipelines. IEEE Software, 38(4), 58–66.
71. Zaharia, M., Xin, R., Armbrust, M., Das, T., Meng, X., Rosen, J., Venkataraman, S., Franklin, M., Ghodsi, A., Gonzalez, J., Shenker, S., & Stoica, I. (2021). Apache Spark: A unified engine for big data processing. Communications of the ACM, 64(11), 114–124.
72. Boehm, M., Kumar, A., Yang, J., & Sze, V. (2023). Feature stores for production machine learning systems. IEEE Data Engineering Bulletin, 46(2), 25–39.
73. Najafabadi, F. A., Bogner, J., Gerostathopoulos, I., & Lago, P. (2024). Architectural components and tools for MLOps systems: A systematic mapping study. Springer Lecture Notes in Computer Science, 14889, 69–85.
74. Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J. F., & Dennison, D. (2021). Hidden technical debt in machine learning systems. Communications of the ACM, 64(7), 54–61.
75. Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., Nagappan, N., Nushi, B., & Zimmermann, T. (2021). Software engineering for machine learning: A case study. Communications of the ACM, 64(7), 56–65.