How to automatically find the batch size when using Accelerate with FSDP2? [D]
A developer is seeking a method to automatically determine or reduce the batch size when using Hugging Face Accelerate with FSDP2 for multi-GPU training. They currently use auto_find_batch_size=True with SFTTrainer for single-GPU training to handle CUDA OOM errors. The developer wants to know if Accelerate can restart distributed training with a smaller batch size or if an external implementation is needed. They also inquire about alternative multi-GPU training approaches that support automatic batch-size detection/recovery if FSDP2 does not.
This query highlights a current limitation in Hugging Face Accelerate's FSDP2 integration, unlike its single-GPU SFTTrainer, regarding automatic batch size adjustment for OOM errors.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 14, 2026, 21:00 UTC
- Ingested
- Sep 14, 2026, 21:00
- Source type
- Dev community
Full text isn't available here.
Read at source →