Edge-to-Cloud Integrated Pipeline Framework for Real-Time Data Processing and Intelligent Workflow Orchestration
Keywords:
Edge computing, cloud computing, real-time data processing, distributed systems, workflow orchestration, stream processing, intelligent scheduling, Kubernetes, Apache Flink, anomaly detection, pipeline automation, resource optimization.Abstract
The rapid growth in the real-time data generated by Internet of Things (IoT) devices, smart infrastructures, and distributed applications poses the need to process the collected data at large scale and latency. Traditional cloud-based systems have a high amount of latency and bandwidth as well as a lack of flexibility to handle scale-on-demand. The present paper introduces an edge-to-cloud integrated pipeline system that is meant to facilitate real-time data processing as well as the orchestration of intelligent workflows. The suggested structure uses edge-level data pre-processing, distributed stream processing based on Apache Flink, and cloud-based aggregation, which are coordinated with the smart orchestration layer. To maximize the latency, throughput, and resources usage, a hybrid scheduling mechanism is built that has a combination of rule-based policy of decision making and adaptive feedback-based strategies. Kubernetes is used to implement the system using container orchestration and prometheus monitoring and observability. The large-scale experimental analysis shows that the framework proposed leads to processing time being shortened by up to 35-fold and throughput increased by about 28-fold with significantly faster fault recovery performance than traditional cloud-only pipelines. The findings indicate the success of the combination of edge and cloud resources with smart orchestration to achieve scalable and robust real-time data processing. The suggested framework is appropriate in next generation distributed applications that need high performance, flexibility and reliability.