Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Coursera

Data Prep with Rust

Pragmatic AI Labs via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Learn to clean, validate, and transform data for machine learning using Agentic tools with Polars, a fast, memory-efficient DataFrame library in Rust. You will take a raw, messy dataset and turn it into a documented, verified artifact: detecting missing values, inferring and checking schemas, catching drift between batches, applying declarative transformations, and building a bronze, silver and gold medallion pipeline that ends in Parquet. Along the way you practice the habit that makes agentic development safe: the agent proposes, the compiler checks that it builds, and you verify the result by running it. Designed for learners new to Rust; no prior Rust experience required.

Syllabus

  • Module 1: Orient — Why Rust + Polars, First Small Win
    • You will see why Rust and Polars suit data preparation, the foundational step where many machine learning projects run into trouble, and what makes development agentic: an AI that works through multiple steps, uses tools and corrects itself under your oversight. You will then drive Claude Code in the terminal, watch a broad request go wrong, and get a first small win from a Rust command-line tool that reports missing values and drops a fully empty column. That sets the habit the rest of the course depends on: the agent proposes, the compiler checks that it builds, and you verify the result by running it.
  • Module 2: Build — The Pipeline
    • You direct an AI coding agent, held to test-driven development by a CLAUDE.md file, to grow one Rust and Polars command-line tool with schema, drift and transform subcommands, each run against a real wine ratings dataset. You infer and check schemas, build a drift baseline of value ranges and categories, and clean the data both in code and through a declarative YAML file. Every check is made to fail on purpose before it is trusted, because catching a wrong type or a drifted value early is far cheaper than finding it after it has broken something downstream.
  • Module 3: Building the Medallion Pipeline
    • You learn the three layers of the medallion architecture and the promise each one makes, then build them as three Rust command-line scripts: bronze lands the raw wine ratings in SQLite, silver cleans them, and gold answers a business question and exports CSV and JSON. You finish by directing an agent to add Parquet output and a check subcommand, because separate layers and verified output are what let you find where a wrong result came from.
  • Module 4: Finish
    • You review the finished Rust data tool, its tests for summary, schema and drift, and the guardrails that let an agent change it safely. You then plan next steps, starting with GitHub Actions that run cargo build and cargo test on every change, because automation keeps confidence high when an agent is making the changes and you remain in charge.

Taught by

Alfredo Deza

Reviews

Start your review of Data Prep with Rust

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.