AIINFRA 200: Production Inference Serving & GPU Orchestration — Syllabus

Course description

This course prepares students to deploy, tune, and operate production-grade large language model inference systems at scale. Students move beyond single-GPU prototyping tools to master industry-standard serving engines (vLLM, TGI, TensorRT-LLM, SGLang) and Kubernetes-native GPU orchestration (GPU Operator, Dynamic Resource Allocation, MIG, KAI Scheduler). By the end of the course, students will be able to benchmark, autoscale, and cost-optimize multi-GPU inference deployments using industry-standard observability tooling.

Contact hours

54 contact hours (3-unit equivalent), delivered over a 16-week term at approximately 3.4 hours per week. This is a non-credit course on the CDCP Certificate of Completion pathway.

Prerequisite

AIINFRA 101 (Docker, Kubernetes, and cloud fundamentals).

Course Learning Outcomes (CLOs)

  1. CLO1: Deploy containerized Linux services and distinguish dev/prototyping tools from production-grade inference stacks.
  2. CLO2: Configure and tune production inference servers (vLLM, TGI, TensorRT-LLM, SGLang) for high-throughput, low-latency serving.
  3. CLO3: Orchestrate GPU resources on Kubernetes using the GPU Operator, Dynamic Resource Allocation, MIG, and the KAI Scheduler.
  4. CLO4: Implement multi-GPU tensor and pipeline parallelism with autoscaling and load balancing for large models.
  5. CLO5: Benchmark, monitor, and cost-optimize inference deployments using Prometheus, Grafana, and disaggregated serving architectures.

Weekly schedule

Week Topic
01 Linux Server Foundations for AI Workloads
02 Docker, GPU Runtime, and Dev-Only Tools (Ollama, LM Studio)
03 Production Inference Fundamentals and the vLLM Engine
04 PagedAttention and Continuous Batching Deep Dive
05 Hugging Face TGI and SGLang Serving Engines
06 TensorRT-LLM Compilation and Optimization
07 NVIDIA Triton Inference Server and Model Repositories
08 Serving Engine Benchmarking Lab (Midterm)
09 Ray Serve and KServe for Scalable Model Serving
10 GPU Orchestration on Kubernetes with the NVIDIA GPU Operator
11 Dynamic Resource Allocation and MIG Partitioning
12 Multi-GPU Tensor and Pipeline Parallelism
13 Autoscaling, Load Balancing, and the KAI Scheduler
14 Disaggregated Prefill/Decode with NVIDIA Dynamo
15 Observability, Cost Analysis, and Right-Sizing
16 Capstone Project & Course Review (Capstone)

Grading

Component Weight
Labs 40%
Discussions 10%
Weekly Quizzes 15%
Midterm 15%
Final Capstone 20%

This course is graded Credit/No-Credit. A minimum of 70% overall is required to pass and earn credit toward the Certificate of Completion.

Course policies

Academic integrity: Students are expected to submit their own work for all labs, quizzes, and the capstone project. Collaboration on concepts is encouraged, but submitted configurations, code, and benchmark results must reflect the student's own effort. Plagiarism or submitting another student's work as your own may result in a failing grade for the assignment and referral to the college's academic integrity process. Late work: Assignments are due as posted in Canvas. Late submissions are accepted up to 7 calendar days after the due date with no penalty, in recognition of the varied schedules of non-credit students; work submitted after 7 days may not be accepted without prior instructor approval. Responsible use of AI: Students are encouraged to use AI coding assistants (e.g., Claude, ChatGPT, Copilot) as learning aids for debugging and exploring configuration options, consistent with real-world infrastructure practice. Students must be able to explain and defend any AI-assisted work they submit, and final architectural decisions, benchmarking analysis, and written reflections must be the student's own. Accessibility: This course is committed to full inclusion of students with disabilities. Students needing accommodations should contact the campus Disabled Students Programs and Services (DSPS) office as early in the term as possible. Course materials are designed to meet accessibility standards; contact the instructor if any content is not accessible to you.

Tools & materials

All tools and materials used in this course are free or free-tier.


Week 01 · Linux Server Foundations for AI Workloads

Notion ID: 392c08fd-0278-8168-b62f-fcd0e9d286c2 Notion URL: https://app.notion.com/p/392c08fd02788168b62ffcd0e9d286c2