How AI Models Scale Beyond a Single GPU Across LLM Workloads
The video "How AI Models Scale Beyond a Single GPU Across LLM Workloads" explains distributed inference for AI models. It covers scaling LLM traffic with Data Parallelism, splitting AI models with Pipeline Parallelism, and dividing LLM layers with Tensor Parallelism. The content also discusses scaling Mixture-of-Experts across GPUs, separating LLM Prefill and Decode Workloads, and combining GPU Parallelism for production AI, concluding with orchestrating distributed LLM Inference.
This video provides a comprehensive overview of distributed inference techniques, detailing multiple parallelism strategies like data, pipeline, and tensor parallelism, unlike many resources that focus on a single method.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 6, 2026, 17:00 UTC
- Ingested
- Oct 6, 2026, 17:00
- Source type
- Unclassified
Full text isn't available here.
Read at source →