跳到正文
YYouTube·
暂不在当前实时榜单

Serve Your Own LLM: vLLM & SGLang, End-to-End

AI 摘要

This course, "Serve Your Own LLM: vLLM & SGLang, End-to-End," offers 12 live lectures starting in November 2026. It covers tracing requests from chat to GPU, building latency and throughput benchmarks, and tuning memory, batching, prefix caching, quantization, multi-GPU, and mixture-of-experts serving. The curriculum also includes production aspects like routers, Kubernetes, metrics, and cost per token, culminating in a capstone project to compare vLLM and SGLang.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年10月8日 11:00 UTC

收录
2026年10月8日 11:00
来源类型
未分类

本站未收录正文。

前往源站阅读 →
来源·YouTube·youtube.com