Skip to content
YYouTube·
Not on the current live radar

Serve Your Own LLM: vLLM & SGLang, End-to-End

AI summary

This course, "Serve Your Own LLM: vLLM & SGLang, End-to-End," offers 12 live lectures starting in November 2026. It covers tracing requests from chat to GPU, building latency and throughput benchmarks, and tuning memory, batching, prefix caching, quantization, multi-GPU, and mixture-of-experts serving. The curriculum also includes production aspects like routers, Kubernetes, metrics, and cost per token, culminating in a capstone project to compare vLLM and SGLang.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Oct 8, 2026, 11:00 UTC

Ingested
Oct 8, 2026, 11:00
Source type
Unclassified

Full text isn't available here.

Read at source →
Source·YouTube·youtube.com