How are you guys getting youtube training data at scale ?[D]
A developer is seeking methods to obtain YouTube video transcripts at scale for a training corpus, specifically for educational content. They initially used the youtube-transcript-api Python library, which worked for a few hundred videos before YouTube began returning empty responses without errors or captchas. Rotating IPs with residential proxies provided a temporary solution, but the issue recurred. The developer is looking for a robust solution that can handle thousands of requests without constant monitoring.
This report highlights a practical challenge in large-scale data collection from YouTube, unlike theoretical discussions, by detailing the failure of common scraping methods after a few hundred requests.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 28, 2026, 18:00 UTC
- Ingested
- Sep 28, 2026, 18:00
- Source type
- Dev community
Full text isn't available here.
Read at source →