跳到正文
RCreddit.com·
暂不在当前实时榜单

How are you guys getting youtube training data at scale ?[D]

AI 摘要

A developer is seeking methods to obtain YouTube video transcripts at scale for a training corpus, specifically for educational content. They initially used the youtube-transcript-api Python library, which worked for a few hundred videos before YouTube began returning empty responses without errors or captchas. Rotating IPs with residential proxies provided a temporary solution, but the issue recurred. The developer is looking for a robust solution that can handle thousands of requests without constant monitoring.

为什么是这条

This report highlights a practical challenge in large-scale data collection from YouTube, unlike theoretical discussions, by detailing the failure of common scraping methods after a few hundred requests.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月28日 18:00 UTC

收录
2026年9月28日 18:00
来源类型
开发者社区

本站未收录正文。

前往源站阅读 →
来源·reddit.com