Develop practical CUDA programming skills using C++ and GPU-based parallel computing. This intermediate path begins with device memory allocation, CPU–GPU data transfers, block configuration, and one-dimensional kernels.
You will progress to two-dimensional grids, row-major indexing, grid-stride loops, and matrix multiplication. You will then use shared memory and thread synchronization to implement tiled algorithms and compare their performance with naive approaches.
Finally, you will apply these techniques to image buffers by mapping pixels to threads, handling boundaries, and building a processing pipeline for grayscale conversion and Sobel edge detection. This path is intended for learners with prior C++ programming experience who want to build and optimize CUDA kernels.