AIINFRA 200: Production Inference Serving & GPU Orchestration — Outcomes & Rubrics

Course Learning Outcomes

By the end of AIINFRA 200, students will be able to:

  1. CLO1 — Deploy containerized Linux services and distinguish dev/prototyping tools from production-grade inference stacks.
  2. CLO2 — Configure and tune production inference servers (vLLM, TGI, TensorRT-LLM, SGLang) for high-throughput, low-latency serving.
  3. CLO3 — Orchestrate GPU resources on Kubernetes using the GPU Operator, Dynamic Resource Allocation, MIG, and the KAI Scheduler.
  4. CLO4 — Implement multi-GPU tensor and pipeline parallelism with autoscaling and load balancing for large models.
  5. CLO5 — Benchmark, monitor, and cost-optimize inference deployments using Prometheus, Grafana, and disaggregated serving architectures.

Together, CLO1–CLO5 map to the certificate's program-level outcomes by building the applied AI infrastructure and architecture competencies — containerization, production model serving, GPU resource orchestration, distributed scaling, and cost/performance optimization — that California employers expect of job-ready AI infrastructure practitioners.

Rubrics