Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Large-scale recommendation systems depend on fast, accurate candidate generation before any ranking takes place. This course teaches you to architect and optimize production-grade retrieval infrastructures capable of surfacing relevant candidates from multi-million item catalogs within strict sub-50ms latency constraints.
You will implement distributed matrix factorization models and Bayesian Personalized Ranking for implicit feedback datasets, build hybrid semantic embedding pipelines to resolve cold-start bottlenecks, and construct two-tower retrieval models optimized for online serving. You will also implement graph-based multi-hop networks using PyTorch Geometric and sequence-based causal transformers for next-item prediction, all executed programmatically within Azure Databricks and Azure ML SDK v2 environments.
This course is designed for Advanced Machine Learning Engineers, Infrastructure Engineers, and Data Platform Architects who are responsible for scaling retrieval services across large item catalogs. You should be comfortable with Python, PyTorch, and core ML evaluation metrics such as Recall@K and NDCG@K.
You will also need access to an active Azure subscription with a configured Azure Databricks workspace and a running compute cluster, as hands-on activities throughout the course are performed in Databricks notebooks. Familiarity with FAISS for vector indexing and similarity search is also expected before beginning.