AIINFRA 200: Production Inference Serving & GPU Orchestration — Outcomes & Rubrics
Course Learning Outcomes
By the end of AIINFRA 200, students will be able to:
- CLO1 — Deploy containerized Linux services and distinguish dev/prototyping tools from production-grade inference stacks.
- CLO2 — Configure and tune production inference servers (vLLM, TGI, TensorRT-LLM, SGLang) for high-throughput, low-latency serving.
- CLO3 — Orchestrate GPU resources on Kubernetes using the GPU Operator, Dynamic Resource Allocation, MIG, and the KAI Scheduler.
- CLO4 — Implement multi-GPU tensor and pipeline parallelism with autoscaling and load balancing for large models.
- CLO5 — Benchmark, monitor, and cost-optimize inference deployments using Prometheus, Grafana, and disaggregated serving architectures.
Together, CLO1–CLO5 map to the certificate's program-level outcomes by building the applied AI infrastructure and architecture competencies — containerization, production model serving, GPU resource orchestration, distributed scaling, and cost/performance optimization — that California employers expect of job-ready AI infrastructure practitioners.
Rubrics