Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Coursera

Big Data on Kubernetes

Packt via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One plan covers every Professional Certificate on Coursera.
Unlock All Certificates
This course provides a comprehensive exploration of deploying and managing big data workloads on Kubernetes. You will learn how containerization and orchestration simplify scaling and managing complex data pipelines in modern enterprise environments. Through hands-on exercises with Apache Spark, Airflow, and Kafka, learners will gain practical skills in building, deploying, and maintaining big data workflows. These real-world applications ensure that you can translate theoretical concepts into production-ready pipelines. Unlike traditional courses, this program combines in-depth architectural insights with actionable implementation guidance, bridging the gap between theory and operational excellence. You will gain exposure to the full stack from data ingestion to consumption. Ideal for data engineers, cloud architects, and developers with basic familiarity with Kubernetes and data processing concepts, this course prepares you to handle production-scale big data applications efficiently.

Syllabus

  • Getting Started with Containers
    • This module introduces the fundamentals of container technology, focusing on Docker installation, setup, and usage for scalable data applications. Learners will gain hands-on experience running containers, deploying applications in different programming languages, and building a simple API service using FastAPI.
  • Kubernetes Architecture
    • This module introduces the foundational components of Kubernetes clusters, including control plane and worker nodes, and explores essential resources such as pods, deployments, services, StatefulSets, ConfigMaps, and Secrets. Learners will gain practical knowledge of how these elements interact to manage and expose applications, as well as how to handle configuration and external traffic using the Gateway API.
  • Kubernetes - Hands On
    • This module guides learners through deploying and managing Kubernetes clusters both locally and in the cloud using AWS EKS and Google Cloud GKE. Participants will gain practical experience containerizing applications, orchestrating deployments, and running APIs on Kubernetes platforms.
  • The Modern Data Stack
    • This module introduces learners to contemporary data architectures, focusing on the Lambda and Kappa patterns for real-time and batch processing. You will explore the evolution of the data lakehouse, compare architectural approaches, and examine technologies for batch and real-time data serving. By the end, you'll understand how modern data stacks enable scalable, flexible analytics.
  • Big Data Processing with Apache Spark
    • This module introduces learners to the essentials of processing large datasets using Apache Spark. You will set up a local PySpark environment, perform data transformations with DataFrames, and utilize Spark SQL for scalable analytics. Practical exercises include working with real-world datasets and understanding key concepts like narrow versus wide transformations and join strategies.
  • Apache Airflow for Building Pipelines
    • This module guides learners through installing Apache Airflow with Docker, building and managing data pipelines using Directed Acyclic Graphs (DAGs), and integrating Airflow with external tools like PostgreSQL and Amazon S3. Learners will also explore different Airflow executors and gain hands-on experience orchestrating real-world data workflows.
  • Apache Kafka for Real-Time Events and Data Ingestion
    • This module introduces learners to the fundamentals of Apache Kafka, including its distributed architecture, data delivery semantics, and integration with real-time processing frameworks like Spark. Participants will gain hands-on experience setting up Kafka, connecting it to databases, and building real-time data pipelines for ingesting and processing streaming data.
  • Deploying the Big Data Stack on Kubernetes
    • This module guides learners through deploying essential big data tools—Spark, Airflow, and Kafka—on Kubernetes using operators and Helm charts. Participants will gain hands-on experience configuring these technologies for scalable data pipelines and managing their deployment in a cloud-native environment.
  • Data Consumption Layer
    • This module guides learners through deploying and configuring Trino and Elasticsearch on Kubernetes to enable efficient querying and real-time data analysis. Participants will gain hands-on experience with distributed SQL engines, data visualization using Kibana, and connecting tools like DBeaver for interactive exploration. The module also covers best practices for managing both structured and unstructured data in modern cloud environments.
  • Building a Big Data Pipeline on Kubernetes
    • This module guides learners through deploying and orchestrating big data tools on Kubernetes, integrating batch and real-time processing systems, and automating data pipelines using Python, SQL, and APIs. Learners will gain hands-on experience configuring Spark jobs, setting up AWS Glue crawlers, and connecting Kafka with Elasticsearch for scalable data workflows.
  • AI/ML Workloads on Kubernetes
    • This module guides learners through deploying and managing generative AI applications on Kubernetes, utilizing tools like Amazon Bedrock, Streamlit, and Retrieval-Augmented Generation (RAG) systems. Participants will address challenges such as bias and limitations in AI models, build and deploy applications, and integrate databases like DynamoDB for scalable solutions.
  • Where to Go from Here
    • This module guides learners through advanced operational considerations for running big data workloads on Kubernetes, including monitoring, security, CI/CD, and automated scalability. It also highlights the essential team skills and cost control strategies needed for successful production deployments. By the end, you'll be equipped to optimize and manage big data environments effectively.

Taught by

Packt - Course Instructors

Reviews

Start your review of Big Data on Kubernetes

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.