Skip to content
RCreddit.com·
Not on the current live radar

How to automatically find the batch size when using Accelerate with FSDP2? [D]

AI summary

A developer is seeking a method to automatically determine or reduce the batch size when using Hugging Face Accelerate with FSDP2 for multi-GPU training. They currently use auto_find_batch_size=True with SFTTrainer for single-GPU training to handle CUDA OOM errors. The developer wants to know if Accelerate can restart distributed training with a smaller batch size or if an external implementation is needed. They also inquire about alternative multi-GPU training approaches that support automatic batch-size detection/recovery if FSDP2 does not.

Why this one

This query highlights a current limitation in Hugging Face Accelerate's FSDP2 integration, unlike its single-GPU SFTTrainer, regarding automatic batch size adjustment for OOM errors.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 14, 2026, 21:00 UTC

Ingested
Sep 14, 2026, 21:00
Source type
Dev community

Full text isn't available here.

Read at source →
Source·reddit.com