Real-Time Intelligent Analytics AI-Driven Data Engineering with PySpark and Microsoft Fabric

Authors

  • Carlos Mendoza Author

Keywords:

Real-time intelligent analytics; analytics pipelines; data engineering; data ingestion; data distribution; PySpark; Microsoft Fabric; data movement; data preparation; data transformation; data enrichment; data management; data serving; data monetization.

Abstract

AI-driven data engineering using PySpark and Microsoft Fabric for real-time intelligent analytics pipelines is synthesized. The study presents an evidence-based review of theoretical foundations and a detailed description of PySpark as a distributed processing engine for data pipelines integrated with Microsoft Fabric, which serves as a design-and-operations fabric, including metadata management. The use of pipelines that accommodate patterns for low latent data processing supports real-time intelligent analytics by working toward zero downtime and latency-driven service delivery.

The demand for near real-time data for analytics is growing rapidly. New disruptive technologies, such as Microsoft Fabric, encompass the entire data movie cycle and enable the combination of different constitutive components of various architectural patterns in a seamless manner. AI and ML powers, using various domains of the AI foundation, are enabling the future of business transformation. Data-engagement pipelines and data-analytics pipelines are critical. The ability to build, manage, and operate these pipelines strikes the right balance between business requirements and technological delivery and cost.

References

1. Ahmed, N., Barczak, A. L. C., Susnjak, T., & Rashid, M. A. (2020). A comprehensive performance analysis of Apache Hadoop and Apache Spark for large scale data sets using HiBench. Journal of Big Data, 7, 110.

2. Armbrust, M., Das, T., Sun, L., Yavuz, B., Zhu, S., Murthy, M., Torres, J., van Hovell, H., Ionescu, A., Łuszczak, A., Switakowski, M., Szafrański, M., Li, X., Ueshin, T., Mokhtar, M., Boncz, P., Ghodsi, A., Paranjpye, S., Senster, P., & Zaharia, M. (2020). Delta Lake: High-performance ACID table storage over cloud object stores. Proceedings of the VLDB Endowment, 13(12), 3411–3424.

3. Mehmood, E., & Anees, T. (2020). Challenges and solutions for processing real-time big data stream: A systematic literature review. IEEE Access, 8, 119123–119143.

4. van Dongen, G., & Van den Poel, D. (2020). Evaluation of stream processing frameworks. IEEE Transactions on Parallel and Distributed Systems, 31(8), 1845–1858.

5. Stahl, F. T., Roesch, E. B., & Dubuc, T. (2021). Mapping the big data landscape: Technologies, platforms and paradigms for real-time analytics of data streams. IEEE Access, 9, 15351–15374.

6. Shah, S., Amannejad, Y., & Krishnamurthy, D. (2021). Diaspore: Diagnosing performance interference in Apache Spark. IEEE Access, 9, 103230–103243.

7. Amistapuram, K. (2023). Privacy-Preserving Machine Learning Models for Sensitive Customer Data in Insurance Systems. Educational Administration: Theory and Practice, 29(4), 5950-5958.

8. Ahmed, N., Barczak, A. L. C., Rashid, M. A., & Susnjak, T. (2021). A parallelization model for performance characterization of Spark big data jobs on Hadoop clusters. Journal of Big Data, 8, 107.

9. Zainab, A., Ghrayeb, A., Abu-Rub, H., Refaat, S. S., & Bouhali, O. (2021). Distributed tree-based machine learning for short-term load forecasting with Apache Spark. IEEE Access, 9, 57372–57384.

10. Taleb, I., Serhani, M. A., Bouhaddioui, C., & Dssouli, R. (2021). Big data quality framework: A holistic approach to continuous quality management. Journal of Big Data, 8, 76.

11. Renggli, C., Rimanic, L., Gurel, N. M., Karlas, B., Wu, W., & Zhang, C. (2021). A data quality-driven view of MLOps. IEEE Data Engineering Bulletin, 44(1), 11–23.

12. He, X., Zhao, K., & Chu, X. (2021). AutoML: A survey of the state-of-the-art. Knowledge-Based Systems, 212, 106622.

13. Kocaman, V., & Talby, D. (2021). Spark NLP: Natural language understanding at scale. arXiv.

14. Migliorini, S., Belussi, A., Quintarelli, E., & Carra, D. (2021). CoPart: A context-based partitioning technique for big data. Journal of Big Data, 8, 21.

15. Priebe, T., Neumaier, S., & Markus, S. (2022). Data warehouse, data lake, data lakehouse, and data mesh: A guide through analytical data architectures. arXiv.

16. Inala, R. Designing Scalable Technology Architectures for Customer Data in Group Insurance and Investment Platforms.

17. Abdalla, H. B. (2022). A brief survey on big data: Technologies, terminologies and data-intensive applications. Journal of Big Data, 9, 107.

18. Akram, A. W., & Alamgir, Z. (2022). Distributed fuzzy clustering algorithm for mixed-mode data in Apache Spark. Journal of Big Data, 9, 121.

19. Zeidan, A., & Vo, H. T. (2022). Efficient spatial data partitioning for distributed kNN joins. Journal of Big Data, 9, 77.

20. Cakir, A., Akın, Ö., Deniz, H. F., & Yılmaz, A. (2022). Enabling real time big data solutions for manufacturing at scale. Journal of Big Data, 9, 118.

21. Kreuzberger, D., Kühl, N., & Hirschl, S. (2022). Machine learning operations (MLOps): Overview, definition, and architecture. IEEE Access, 10, 73843–73865.

22. Davuluri, P. S. L. (2023). AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems (December 15, 2023).

23. Symeonidis, G., Nerantzis, E., Kazakis, A., & Papakostas, G. A. (2022). MLOps—Definitions, tools and challenges. arXiv.

24. Almeida, R., et al. (2023). Time series big data: A survey on data stream frameworks, analysis and algorithms. Journal of Big Data, 10, 83.

25. Chen, W., Milosevic, Z., Rabhi, F. A., & Berry, A. (2023). Real-time analytics: Concepts, architectures, and ML/AI considerations. IEEE Access, 11, 71634–71657.

26. Xu, G., Song, M., Leng, Z., & Jia, Z. (2023). Simulation research on fast matching of big data based on Spark. IEEE Access.

27. Brito, C., Ferreira, P. G., Portela, B. L., et al. (2023). Privacy-preserving machine learning on Apache Spark. IEEE Access.

28. Mazumdar, D., Hughes, J., & Onofre, J. B. (2023). The data lakehouse: Data warehousing and more. arXiv.

29. Goedegebuure, A., Kumara, I., Driessen, S., Di Nucci, D., Monsieur, G., van den Heuvel, W.-J., & Tamburri, D. A. (2023). Data mesh: A systematic gray literature review. arXiv.

30. Yandamuri, U. S. (2023). An Intelligent Analytics Framework Combining Big Data and Machine Learning for Business Forecasting. Zenodo.

31. Azeem, M., Abualsoud, B. M., & Priyadarshana, D. (2023). Mobile big data analytics using deep learning and Apache Spark. Mesopotamian Journal of Big Data.

32. Hasan, Z., Xing, H. J., & Magray, M. I. (2022). Big data machine learning using Apache Spark MLlib. Mesopotamian Journal of Big Data.

33. Alvarez Rodriguez, S., Chakraborty, J., Chu, A., Jimenez, I., LeFevre, J., Maltzahn, C., & Uta, A. (2021). Zero-cost, Arrow-enabled data interface for Apache Spark. arXiv.

34. Shaikh, S. A., Mariam, K., Kitagawa, H., & Kim, K.-S. (2020). GeoFlink: A distributed and scalable framework for the real-time processing of spatial streams. arXiv.

Additional Files

Published

2023-12-12

Data Availability Statement

none

How to Cite

Real-Time Intelligent Analytics AI-Driven Data Engineering with PySpark and Microsoft Fabric. (2023). European Data Science Journal (EDSJ), 1(01). https://esa-research.org/index.php/EDSJ/article/view/106

Most read articles by the same author(s)

Similar Articles

11-20 of 47

You may also start an advanced similarity search for this article.