This intermediate path covers the design and implementation of an end-to-end data engineering architecture on AWS. You will structure an S3 data lake with raw, processed, and curated zones, ingest streaming data, and catalog datasets for downstream use.
You will build ETL pipelines with AWS Glue and PySpark, convert JSON data to Parquet, and query data with Amazon Athena. You will also create a Redshift Serverless analytics warehouse, model a star schema, and optimize data loading and query performance.
Finally, you will automate workflows with Lambda and EventBridge by responding to S3 uploads, running scheduled processes, and launching Glue jobs. This path is intended for learners with foundational cloud and data skills who want practical experience building AWS data pipelines.