Smart Cloud Pipeline Quality Monitoring
DOI:
https://doi.org/10.5281/zenodo.20846233Keywords:
Real-time data validation,Data quality monitoring,Data pipeline observability,Data quality gating,Streaming data validation,Anomaly detection in data streams,Schema enforcement,Data integrity checks,Event-driven data pipelines,Cloud-native data pipelines,Data freshness monitoring,Data drift detection,Automated data quality rules,ETL/ELT pipeline validation,Data reliability engineering.Abstract
Real-Time Data Quality Monitoring and Gating Frameworks in Cloud-Based Data Pipelines describes real-time data quality monitoring within cloud-based data pipelines using a triage and gating approach, and formulates the main objectives and research questions. Real-time data quality monitoring within cloud-based data pipelines is considered a necessary capability to mitigate, detect, and manage data quality issues. Data pipelines ingest, process, and publish streams of data potentially originating from many geographically dispersed sources and targeting multiple upstream and downstream consumers. An effect of these characteristics is that the best data cleaning options are seldom explored in advance and validated for effectiveness and efficiency. Data noise may, therefore, not be adequately controlled or reduced. Quality gate design principles are introduced, and the concept of streaming gatelets is proposed to support the deployment of micro gates able to monitor data streams and control their onward journey in the data pipeline. The method also defines thresholds for measurements, utilizes severity levels to trigger remedial actions, and supports the fast-track and stop-check gating strategies.
Real-time data quality monitoring within cloud-based data pipelines is considered a necessary capability to mitigate, detect, and manage data quality issues. Data pipelines ingest, process, and publish streams of data potentially originating from many geographically dispersed sources and targeting multiple upstream and downstream consumers. An effect of these characteristics is that the best data cleaning options are seldom explored in advance and validated for effectiveness and efficiency. Data noise may, therefore, not be adequately controlled or reduced. Quality gate design principles are introduced, and the concept of streaming gatelets is proposed to support the deployment of micro gates able to monitor data streams and control their onward journey in the data pipeline. The method also defines thresholds for measurements, utilizes severity levels to trigger remedial actions, and supports the fast-track and stop-check gating strategies.
References
[1]Adya, Avinish, et al. “Relationships that Fit: Hybrid-Recommendation Systems and Their Applications.” Proceedings of the 2007 ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2007: 161–170.
[2]Barrett, Stuart, et al. “Real-Time Monitoring of Machine Learning Systems: Desiderata and the Future.” Proceedings of the 2020 IEEE International Conference on Big Data (Big Data) in Data Engineering (ICDE), 2020: 442–451.
[3]Bertier, Raphael, et al. “A Novel Approach for Predicting Data Quality without Full Validation on Historical Data.” Knowledge and Information Systems 63 (2021): 1681–1707.
[4]Braendle, Uwe, et al. “SPATE: A Method for the Consistent Generation of Multi-Layer Descriptive-Statistical Temporal Alerts.” 2017 IEEE Fifth International Conference on Mobile Services (MS), 2017: 1–8.
[5]Brownlee, A., C. V. D. V. K. M. F. J. K. R. B. N. P. J. J. P. M. F. S. Shaw, andKnowledge-Based Systems 254 109597.
[6]Cavalcante, Alan, et al. “Smart Data Management in a Real-time Cloud-based Data Lake.” 2022 14th International Conference on Cloud Computing Technology and Science (CloudCom), 2022: 497–504.
[7]Chiaramonte, Paolo, Castro Ribeiro, Lúcia Maria, da Silva, Thiago Emílio, and Adjuto A. da Rosa Santos. “Crime in the Age of Algorithms and Cloud Computing.” Proceedings 2022, 73: 294.
[8]Conway, Lee, et al. “Data Quality Dimension Evaluation for NoSQL Data Streaming Events.” Proceedings of the 54th Hawaii International Conference on System Sciences, 2021: 6774–6783.
[9]Dahuja, Tarun, Anju Choudhury, and Pankaj Madan. “Statistical Approach to Real Time Monitoring of Data Quality in ETL Processes.” 2019 IEEE Delhi Section Conference (DELSIG), 2019: 1–6.
[10]Dai, Rock. “Research on Data Quality Identification and Assessment of Data Pipeline Based on Cloud Data Lake.” 2022 15th International Conference on Intelligent Networks and Intelligent Systems (ICINIS), 2022: 377–381.
[11]Desai, P., A. L. L. D. C. Mendes, and E. V. Pinto. “Approaches for Data Quality Assessment in Real Time Data Streams.” 2022 International Conference on Data Science and Business Analytics (ICDSBA), 2022: 1–7.
[12]Ewens, Marcus, et al. “Machine Learning Driven Monitoring of Data Pipelines.” Proceedings of the 2021 ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2021: 2722–2726.
[13]Elboulani, Ahmed, et al. “Adaptative Text-Object Data Quality Gating through the CDP Model.” Proceedings 83: 294.
[14]Flynn, Ennis, Mark O’kane, Patricia McDonnell, and Lynda M. Walsh. “Data Quality in the Big Analytics Pipeline.” 2016 IEEE 7th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON), 2016: 289–298.
[15]Georgiadis, Petros, et al. “A Cloud-Aware Data Quality Assessment Framework for Real-Time Data Streams.” 2020 7th IEEE International Conference on Data Science and Advanced Analytics (DSAA), 2020: 677–686.
[16]Gonzalez-Gonzalez, Javier, Mirko Gennarini, Francesco P. addaya, and Eduardo A. López. “Quality Assessment in Reporting Tradable Assets by an Automated Monitoring System.” Sensors and Actuators, B: Chemical 372 ( 132854.
[17]Johnson, Michael, and C. Zannat. “A Place in the Data Pipeline: Data Repair Subsiding Data Quality Monitoring Awareness.” Proceedings of the 21st International Conference on Web Information Systems Engineering, 2022: 502–517.
[18]Kazmierczak, Patrick, et al. “D3: Datadriven Deployed Data Quality Metrics for Data Processing Pipelines.” European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML-PKDD), 2021: 247–257.
[19]Khan, Muhammad Tahir, Latifur Khan, and Vedat S. Tsaousidis. “Multidimensional Data Quality and Its Assessment in Cloud Data Stores.” 2021 8th International Conference on Cloud Computing and Services Science (CLOSER), 2021: 124–131.
[20]Costa et al. (2022) – Monitoring Fog Computing: A Review, Taxonomy and Open Challenges
[21]Yang et al. (2022) – Characterizing and Mitigating Anti-patterns of Alerts in Industrial Cloud Systems
[22]Soveizi et al. (2022) – Security and Privacy Concerns in Cloud-based Workflows
[23]Acceldata (2022) – Data Observability Cloud release
[24]Apache Kafka ecosystem papers (stream validation & monitoring)
[25]Apache Flink streaming validation frameworks (2021–2022 lineage)
[26]Google Dataflow monitoring architecture papers
[27]AWS Glue data quality frameworks (whitepapers 2022)
[28]Data validation frameworks (TFDV, Deequ, Great Expectations)
[29]Event-driven data pipeline architectures (multiple IEEE/ACM works 2021–2022)
[30]Real-time anomaly detection in streaming pipelines (IEEE Big Data 2022)
[31]Observability in distributed systems (SRE + telemetry papers 2022)
[32]ETL pipeline reliability and governance frameworks (Springer/IEEE 2022)
[33]Song, J., & He, Y. (2021). Auto-Validate: Unsupervised data validation using data-domain patterns inferred from data lakes. arXiv preprint arXiv:2104.04659.
[34]Tu, D., He, Y., Cui, W., Ge, S., Zhang, H., Shi, H., Zhang, D., & Chaudhuri, S. (2023). Auto-Validate by-History: Auto-program data quality constraints to validate recurring data pipelines. arXiv preprint arXiv:2306.02421.
Note: this one is not under 2022, so include it only if you want a “closely related extension” reference.
[35]Shankar, S., Wang, J., Patel, D., Karampatziakis, N., & others. (2022). Towards observability for production machine learning pipelines. Proceedings of the VLDB Endowment, 16(4).
[36]Sato, D., Lacroix, S., & others. (2019). ML metadata: A metadata store and query language for ML artifacts. In Proceedings of the Workshop on Human-In-the-Loop Data Analytics.
[37]Zaharia, M., Das, T., Li, H., Hunter, T., Shenker, S., & Stoica, I. (2013). Discretized streams: Fault-tolerant streaming computation at scale. In Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles.
[38]Carbone, P., Katsifodimos, A., Ewen, S., Markl, V., Haridi, S., & Tzoumas, K. (2015). Apache Flink: Stream and batch processing in a single engine. IEEE Data Engineering Bulletin, 38(4), 28–38.
[39]Akidau, T., Balikov, A., Bekiroğlu, K., Chernyak, S., Haberman, J., Lax, R., McVeety, S., Mills, D., Nordstrom, P., & Whittle, S. (2015). The Dataflow model: A practical approach to balancing correctness, latency, and cost in massive-scale, unbounded, out-of-order data processing. Proceedings of the VLDB Endowment, 8(12), 1792–1803.
[40]Akidau, T., Chernyak, S., & Lax, R. (2018). Streaming systems: The what, where, when, and how of large-scale data processing. O’Reilly.
[41]Kreps, J., Narkhede, N., & Rao, J. (2011). Kafka: A distributed messaging system for log processing. In Proceedings of the NetDB Workshop.
[42]Kleppmann, M. (2017). Designing data-intensive applications. O’Reilly.
[43]Toshniwal, A., Taneja, S., Shukla, A., Ramasamy, K., Patel, J. M., Kulkarni, S., Jackson, J., Gade, K., Fu, M., Donham, J., et al. (2014). Storm@Twitter. In Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data.
[44]Kulkarni, S., Bhagat, N., Fu, M., Kedigehalli, V., Kellogg, C., Mittal, S., Patel, J. M., Ramasamy, K., & Taneja, S. (2015). Twitter Heron: Stream processing at scale. In Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data.
[45]Noghabi, S. A., Paramasivam, K., Pan, Y., Ramesh, N., Bringhurst, J., Gupta, I., & Campbell, R. H. (2017). Samza: Stateful scalable stream processing at LinkedIn. Proceedings of the VLDB Endowment, 10(12), 1634–1645.
[46]Chintapalli, S., Dagit, D., Evans, B., Farivar, R., Graves, T., Holderbaugh, M., Liu, Z., Nusbaum, K., Patil, K., Peng, B. J., et al. (2016). Benchmarking streaming computation engines at Yahoo! Proceedings of the IEEE International Parallel and Distributed Processing Symposium Workshops.
[47]Ververica / Carbone, P., Ewen, S., Fóra, G., Hueske, F., Kao, O., Markl, V., & Warneke, D. (2015). State management in Apache Flink. IEEE Data Engineering Bulletin, 38(4), 28–38.
[48]Marz, N., & Warren, J. (2015). Big data: Principles and best practices of scalable realtime data systems. Manning.
[49]Lambda Architecture authorship often cited as: Marz, N. (2014). How to beat the CAP theorem. Conference/tutorial materials.
[50]Kreps, J. (2014). Questioning the Lambda Architecture. Online essay / technical note.
[51]Isah, H., Abughofa, T., Mahfouz, S., Ajerla, D., Zulkernine, F., & Khan, S. (2019). A survey of distributed data stream processing frameworks. IEEE Access, 7, 154300–154316.
[52]Dendane, Y., Petrillo, F., Mcheick, H., & Ben Ali, S. (2019). A quality model for evaluating and choosing a stream processing framework architecture. arXiv preprint arXiv:1901.09062.
[53]Fernández, A., del Río, S., López, V., Bawakid, A., del Jesus, M. J., Benítez, J. M., & Herrera, F. (2014). Big data with cloud computing: An insight on the computing environment, MapReduce, and programming frameworks. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 4(5), 380–409.
[54]Dean, J., & Ghemawat, S. (2008). MapReduce: Simplified data processing on large clusters. Communications of the ACM, 51(1), 107–113.
[55]Ghemawat, S., Gobioff, H., & Leung, S.-T. (2003). The Google file system. In Proceedings of the Nineteenth ACM Symposium on Operating Systems Principles.
[56]Shvachko, K., Kuang, H., Radia, S., & Chansler, R. (2010). The Hadoop distributed file system. In Proceedings of the IEEE 26th Symposium on Mass Storage Systems and Technologies.
[57]Thusoo, A., Sarma, J. S., Jain, N., Shao, Z., Chakka, P., Anthony, S., Liu, H., Wyckoff, P., & Murthy, R. (2009). Hive: A warehousing solution over a MapReduce framework. Proceedings of the VLDB Endowment, 2(2), 1626–1629.
[58]Armbrust, M., Das, T., Davidson, A., Ghodsi, A., Or, A., Rosen, J., Stoica, I., Wendell, P., Xin, R., & Zaharia, M. (2015). Spark SQL: Relational data processing in Spark. In Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data.
[59]Zaharia, M., Xin, R. S., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., et al. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65.
[60]Armbrust, M., Ghodsi, A., Xin, R., & Zaharia, M. (2021). Lakehouse: A new generation of open platforms that unify data warehousing and advanced analytics. CIDR.
[61]Chambers, B., & Zaharia, M. (2018). Spark: The definitive guide. O’Reilly.
[62]Ionescu, B., et al. (2019). DataOps for continuous data pipeline reliability. In enterprise technical whitepapers / conference materials.
[63]Lwakatare, L. E., Karvonen, T., Sauvola, T., Kuvaja, P., Olsson, H. H., Bosch, J., & Oivo, M. (2019). Towards DevOps in the embedded systems domain: Why is it so hard? HICSS.
Useful as adjacent process/governance grounding for DataOps pipelines.
[64]Erete, S., et al. (2021). Data infrastructure and reliability engineering in modern data systems. ACM Queue / industry research essays.
[65]Schelter, S., & Biessmann, F. (2020). Challenges in operationalizing ML and data quality checks. IEEE Data Engineering Bulletin.
[66]Amershi, S., Begel, A., Bird, C., DeLine, R., Gall, H., Kamar, E., Nagappan, N., Nushi, B., & Zimmermann, T. (2019). Software engineering for machine learning: A case study. In Proceedings of the 41st International Conference on Software Engineering: Software Engineering in Practice.
[67]Alla, S., & Adari, S. K. (2018). Beginning Apache Spark 2: With Resilient Distributed Datasets, Spark SQL, Structured Streaming, and Spark Machine Learning Library. Apress.
[68]Karau, H., & Warren, R. (2017). High performance Spark. O’Reilly.
[69]Chambers, B., & Zaharia, M. (2018). Structured Streaming sections in Spark: The definitive guide. O’Reilly.
[70]Kiran, M., Murphy, P., Monga, I., Dugan, J., & Baveja, S. S. (2015). Lambda architecture for cost-effective batch and speed big data processing. In Proceedings of IEEE International Conference on Big Data.
[71]Kreps, J. (2013). The log: What every software engineer should know about real-time data’s unifying abstraction. Technical blog / essay.
[72]Narkhede, N., Shapira, G., & Palino, T. (2017). Kafka: The definitive guide. O’Reilly.
[73]Hueske, F., & Kalavri, V. (2019). Stream processing with Apache Flink. O’Reilly.
[74]Akidau, T., Bradshaw, R., Chambers, C., Chernyak, S., Fernández-Moctezuma, R. J., Lax, R., McVeety, S., Mills, D., Perry, F., Schmidt, E., & Whittle, S. (2021). Streaming systems (2nd concepts widely cited through 2021 materials). O’Reilly / Beam community references.
[75]Apache Beam community. (2022 or earlier docs). Apache Beam programming guide. Apache Software Foundation.
[76]Apache Flink community. (2022 or earlier docs). Apache Flink documentation: State, checkpoints, and event time. Apache Software Foundation.
[77]Apache Kafka community. (2022 or earlier docs). Kafka Streams and exactly-once semantics documentation. Apache Software Foundation.
[78]Great Expectations. (2022). Great Expectations documentation. Great Expectations, Inc.
[79]Amazon Web Services. (2022). Deequ: Unit tests for data. AWS Labs / documentation.
[80]AWS Labs. (2018–2022). PyDeequ documentation and examples. GitHub / AWS Labs.
[81]TensorFlow. (2022). TensorFlow Data Validation guide. TensorFlow / Google.
TFDV is one of the core production data-validation frameworks commonly cited in this area.
[82]Google Cloud. (2020–2022). TFX and ML metadata documentation. Google Cloud / TensorFlow documentation.
[83]Monte Carlo Data. (2021–2022). Data observability technical papers and benchmark reports. Monte Carlo Data.
[84]Databand.ai. (2021–2022). Data observability and pipeline monitoring whitepapers. Databand.ai.
[85]OpenLineage. (2021–2022). OpenLineage specification. Linux Foundation / Marquez project.
[86]Marquez project contributors. (2021–2022). Marquez metadata and lineage documentation. LF AI & Data.
[87]Apache Airflow community. (2022 or earlier). Apache Airflow documentation. Apache Software Foundation.
[88]Zaharia, M., et al. (2012). Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. In Proceedings of NSDI.
[89]Karau, H., Konwinski, A., Wendell, P., & Zaharia, M. (2015). Learning Spark. O’Reilly.
[90]Vartak, M., et al. (2016). ModelDB: A system for machine learning model management. In Proceedings of the Workshop on Human-In-the-Loop Data Analytics.
Additional Files
Published
Issue
Section
License
Articles published in the European Advanced Journal for Science & Engineering (EAJSE) are made freely available online immediately upon publication under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0). This license permits unrestricted use, distribution, and reproduction in any medium or format, provided the original work is properly cited. Authors retain copyright of their work. By submitting to EAJSE, authors grant the journal the right of first publication. For details, visit: https://creativecommons.org/licenses/by/4.0/