Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Coursera

Streaming Data Pipelines with Kafka and Spark

Edureka via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Build practical expertise in real-time data engineering using Apache Kafka, Apache Spark, Spark Structured Streaming, PySpark, Python, Docker, and Docker Compose. You will learn how data engineers design, process, monitor, and maintain streaming pipelines that continuously move event data from source systems to reliable downstream outputs. You will begin by exploring event streaming fundamentals and the architecture of Apache Kafka. You will examine brokers, topics, partitions, producers, consumers, offsets, consumer groups, message ordering, and delivery behavior. Through guided demonstrations, you will create Kafka topics, publish events with producers, consume messages, and observe how Kafka distributes and manages streaming data. You will then move into stream processing with Spark Structured Streaming, where you will work with streaming DataFrames, schemas, micro-batch execution, transformations, aggregations, and output sinks. You will also explore event-time processing, late-arriving data, stateful operations, checkpointing, and recovery to understand how Spark maintains progress and processes continuously arriving events reliably. Finally, you will focus on streaming data quality, monitoring, failure handling, and reliable delivery. You will validate streaming records, identify malformed or problematic events, monitor pipeline behavior, troubleshoot processing issues, and apply recovery practices for continuous workloads. The course concludes with a reliable open-source streaming pipeline project that brings together Kafka-based event ingestion, Spark processing, quality validation, monitoring, recovery, and dependable output delivery. By the end of this course, you will be able to: - Explain the fundamentals of event streaming and Apache Kafka architecture. - Work with Kafka brokers, topics, partitions, producers, consumers, offsets, and consumer groups. - Explain how partitioning, ordering, and delivery behavior influence streaming pipelines. - Create and manage Kafka-based producer and consumer workflows. - Process continuously arriving events using Spark Structured Streaming. - Apply schemas and transformations to streaming DataFrames. - Perform streaming aggregations and deliver processed results to output sinks. - Work with event-time processing, late-arriving data, and stateful operations. - Apply checkpointing and recovery techniques to maintain processing continuity. - Validate streaming events and handle malformed or unreliable records. - Monitor Kafka and Spark streaming workflows and identify operational issues. - Troubleshoot common failures across streaming data pipelines. - Apply reliability practices for continuous processing and delivery. - Build an end-to-end streaming data pipeline using Kafka and Spark. Designed for data engineers, aspiring streaming data engineers, software developers, data platform professionals, and technical professionals working with real-time data systems, this course prepares you to build scalable, reliable, and maintainable streaming pipelines using modern open-source technologies.

Syllabus

  • Event Streaming with Apache Kafka
    • Explore the fundamentals of event-driven streaming systems using Apache Kafka. Learn how streaming architectures are designed, how events are produced and consumed, and how Kafka handles topics, partitions, consumer groups, retention, and replay to support reliable real-time data processing.
  • Streaming Processing with PySpark
    • Develop practical stream processing skills using PySpark and Structured Streaming. Learn how to process real-time data, work with event time and windows, manage stateful streams, handle late-arriving data, and build reliable streaming pipelines using checkpointing and appropriate output modes.
  • Quality, Monitoring, and Delivery
    • Learn how to build reliable and production-ready streaming data pipelines. Explore data quality, schema validation, monitoring, operational readiness, capacity planning, and resilient pipeline design while developing the skills needed to identify failures and maintain dependable streaming workflows.

Taught by

Edureka

Reviews

Start your review of Streaming Data Pipelines with Kafka and Spark

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.