Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

DataCamp

Scaling and Optimizing Data Pipelines with Polars

via DataCamp

Overview

Take your Polars skills to production scale.

Take your Polars skills to production scale. Learn to read query plans and unlock the optimizer's full potential, work efficiently with Parquet, CSV, and database sources, and exploit advanced dtypes like List, Struct, Categorical, and Enum. You'll also stream large queries to disk, process data in batches, and build testable pipelines with built-in assertions. By the end, you'll be equipped to build high-performing data workflows that handle datasets of any size.

Syllabus

  • Query Optimization Deep Dive
    • Learn how to keep queries lazy for maximum optimization, read and interpret query plans, and unlock fast paths with profiling and sorted data.
  • Efficient Data Input and Output
    • Learn how to read and write Parquet files, parse messy CSVs, scan multifile and hive-partitioned datasets, and query databases from Polars.
  • Advanced Dtypes for Optimal Analysis
    • This chapter covers working with List and Struct columns, encoding repeated strings as Categorical and Enum dtypes, and reducing memory use through numeric downcasting.
  • Working with Polars at Scale
    • Learn how to use the streaming and GPU engines, sink large query results directly to disk with partitioning, and test pipelines with Polars' built-in assertions.

Taught by

Liam Brannigan

Reviews

4.9 rating at DataCamp based on 90 ratings

Start your review of Scaling and Optimizing Data Pipelines with Polars

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.