Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

edX

Introduction to CUDA Kernel Programming

via edX

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates

Develop practical CUDA programming skills using C++ and GPU-based parallel computing. This intermediate path begins with device memory allocation, CPU–GPU data transfers, block configuration, and one-dimensional kernels.

You will progress to two-dimensional grids, row-major indexing, grid-stride loops, and matrix multiplication. You will then use shared memory and thread synchronization to implement tiled algorithms and compare their performance with naive approaches.

Finally, you will apply these techniques to image buffers by mapping pixels to threads, handling boundaries, and building a processing pipeline for grayscale conversion and Sobel edge detection. This path is intended for learners with prior C++ programming experience who want to build and optimize CUDA kernels.

Syllabus

  • Allocate GPU memory and transfer data safely between CPUs and GPUs
  • Configure CUDA blocks and grids for one- and two-dimensional workloads
  • Implement scalable kernels using linear indexing and grid-stride loops
  • Build naive and tiled matrix multiplication kernels
  • Optimize global memory access with shared memory and thread synchronization
  • Create GPU image-processing pipelines for grayscale conversion and edge detection
  • Validate kernel outputs and benchmark optimized implementations

Reviews

Start your review of Introduction to CUDA Kernel Programming

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.