This course prepares students to deploy, tune, and operate production-grade large language model inference systems at scale. Students move beyond single-GPU prototyping tools to master industry-standard serving engines (vLLM, TGI, TensorRT-LLM, SGLang) and Kubernetes-native GPU orchestration (GPU Operator, Dynamic Resource Allocation, MIG, KAI Scheduler). By the end of the course, students will be able to benchmark, autoscale, and cost-optimize multi-GPU inference deployments using industry-standard observability tooling.
54 contact hours (3-unit equivalent), delivered over a 16-week term at approximately 3.4 hours per week. This is a non-credit course on the CDCP Certificate of Completion pathway.
AIINFRA 101 (Docker, Kubernetes, and cloud fundamentals).
| Week | Topic |
|---|---|
| 01 | Linux Server Foundations for AI Workloads |
| 02 | Docker, GPU Runtime, and Dev-Only Tools (Ollama, LM Studio) |
| 03 | Production Inference Fundamentals and the vLLM Engine |
| 04 | PagedAttention and Continuous Batching Deep Dive |
| 05 | Hugging Face TGI and SGLang Serving Engines |
| 06 | TensorRT-LLM Compilation and Optimization |
| 07 | NVIDIA Triton Inference Server and Model Repositories |
| 08 | Serving Engine Benchmarking Lab (Midterm) |
| 09 | Ray Serve and KServe for Scalable Model Serving |
| 10 | GPU Orchestration on Kubernetes with the NVIDIA GPU Operator |
| 11 | Dynamic Resource Allocation and MIG Partitioning |
| 12 | Multi-GPU Tensor and Pipeline Parallelism |
| 13 | Autoscaling, Load Balancing, and the KAI Scheduler |
| 14 | Disaggregated Prefill/Decode with NVIDIA Dynamo |
| 15 | Observability, Cost Analysis, and Right-Sizing |
| 16 | Capstone Project & Course Review (Capstone) |
| Component | Weight |
|---|---|
| Labs | 40% |
| Discussions | 10% |
| Weekly Quizzes | 15% |
| Midterm | 15% |
| Final Capstone | 20% |
This course is graded Credit/No-Credit. A minimum of 70% overall is required to pass and earn credit toward the Certificate of Completion.
Academic integrity: Students are expected to submit their own work for all labs, quizzes, and the capstone project. Collaboration on concepts is encouraged, but submitted configurations, code, and benchmark results must reflect the student's own effort. Plagiarism or submitting another student's work as your own may result in a failing grade for the assignment and referral to the college's academic integrity process. Late work: Assignments are due as posted in Canvas. Late submissions are accepted up to 7 calendar days after the due date with no penalty, in recognition of the varied schedules of non-credit students; work submitted after 7 days may not be accepted without prior instructor approval. Responsible use of AI: Students are encouraged to use AI coding assistants (e.g., Claude, ChatGPT, Copilot) as learning aids for debugging and exploring configuration options, consistent with real-world infrastructure practice. Students must be able to explain and defend any AI-assisted work they submit, and final architectural decisions, benchmarking analysis, and written reflections must be the student's own. Accessibility: This course is committed to full inclusion of students with disabilities. Students needing accommodations should contact the campus Disabled Students Programs and Services (DSPS) office as early in the term as possible. Course materials are designed to meet accessibility standards; contact the instructor if any content is not accessible to you.
All tools and materials used in this course are free or free-tier.
Notion ID: 392c08fd-0278-8168-b62f-fcd0e9d286c2 Notion URL: https://app.notion.com/p/392c08fd02788168b62ffcd0e9d286c2