Get 20% off all career paths from fullstack to AI
Learn the Skills Netflix, Meta, and Capital One Actually Hire For
Overview
Google, IBM & Meta Certificates – 40% Off
One plan covers every Professional Certificate on Coursera.
Unlock All Certificates
Discover how to leverage dask-sql for scalable end-to-end data engineering in this 26-minute PyCon US talk. Learn to overcome the challenges of accessing data trapped in Hive/Spark-based datalakes or complex SQL queries. Explore the capabilities of dask-sql, which enables Python developers to create comprehensive data projects without extensive knowledge of JVM/Hadoop ecosystems. Gain insights into performing SQL data extraction from datalakes and Hive tables using Python and dask-sql. Understand how to refine and utilize extracted data for machine learning, analytics, or transformation workloads with popular PyData tools. Delve into the innovative design of dask-sql, which combines SQL optimization from Apache Calcite, scalable dataframe operations via Dask, and integration with the Hive metastore data catalog. Follow along with a demo and walkthrough of dask-sql, covering topics such as the enterprise data processing pipeline and practical implementation. Access accompanying slides for further reference and study.
Syllabus
Introduction
Enterprise Data Processing Pipeline
DaskSQL
Demo
DaskSQL Walkthrough
Taught by
PyCon US