Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

edX

Data Processing for LLMs

via edX

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates

This beginner-friendly path introduces the core data processing techniques used in NLP and large language model applications. You will learn how raw text is cleaned, structured, tokenized, transformed, and stored so it can be used effectively in search, retrieval, training, and knowledge management workflows. The path covers foundational preprocessing methods such as text cleaning, bag-of-words, TF-IDF, embeddings, and modern tokenization approaches including BPE, WordPiece, and SentencePiece. You will also explore how chunking supports efficient processing of long documents and how vector databases enable structured retrieval. As you progress, you will examine scalable preparation strategies for large datasets, including data collection, deduplication, filtering, augmentation, and efficient storage. The focus is on building reliable pipelines that improve data quality and support downstream LLM performance.

Syllabus

  • Clean and preprocess text for NLP and LLM workflows
  • Apply tokenization methods such as BPE, WordPiece, and SentencePiece
  • Vectorize text using bag-of-words, TF-IDF, and embeddings
  • Chunk long documents for efficient retrieval and LLM processing
  • Store processed text and embeddings for structured search workflows
  • Optimize large-scale datasets through deduplication, filtering, and augmentation

Reviews

Start your review of Data Processing for LLMs

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.