Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

Dask-SQL - Empowering Pythonistas for Scalable End-to-End Data Engineering

PyCon US via YouTube

Overview

Coursera Flash Sale
40% Off Coursera Plus for 3 Months!
Grab it
Discover how to leverage dask-sql for scalable end-to-end data engineering in this 26-minute PyCon US talk. Learn to overcome the challenges of accessing data trapped in Hive/Spark-based datalakes or complex SQL queries. Explore the capabilities of dask-sql, which enables Python developers to create comprehensive data projects without extensive knowledge of JVM/Hadoop ecosystems. Gain insights into performing SQL data extraction from datalakes and Hive tables using Python and dask-sql. Understand how to refine and utilize extracted data for machine learning, analytics, or transformation workloads with popular PyData tools. Delve into the innovative design of dask-sql, which combines SQL optimization from Apache Calcite, scalable dataframe operations via Dask, and integration with the Hive metastore data catalog. Follow along with a demo and walkthrough of dask-sql, covering topics such as the enterprise data processing pipeline and practical implementation. Access accompanying slides for further reference and study.

Syllabus

Introduction
Enterprise Data Processing Pipeline
DaskSQL
Demo
DaskSQL Walkthrough

Taught by

PyCon US

Reviews

Start your review of Dask-SQL - Empowering Pythonistas for Scalable End-to-End Data Engineering

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.