How to automatically find the batch size when using Accelerate with FSDP2? [D]
A developer is seeking a method to automatically determine or reduce the batch size when using Hugging Face Accelerate with FSDP2 for multi-GPU training. They currently use auto_find_batch_size=True with SFTTrainer for single-GPU training to handle CUDA OOM errors. The developer wants to know if Accelerate can restart distributed training with a smaller batch size or if an external implementation is needed. They also inquire about alternative multi-GPU training approaches that support automatic batch-size detection/recovery if FSDP2 does not.
This query highlights a current limitation in Hugging Face Accelerate's FSDP2 integration, unlike its single-GPU SFTTrainer, regarding automatic batch size adjustment for OOM errors.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月14日 21:00 UTC
- 收录
- 2026年9月14日 21:00
- 来源类型
- 开发者社区
本站未收录正文。
前往源站阅读 →