The Open ASR Leaderboard Adds Its First Global South Language
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
官方发布涉及Hugging Face 模型访问、订阅权益规则,适合跟踪产品开放节奏和用户影响。
Hugging Face与Voice Arena合作,为印地语和印度英语推出了开放式自动语音识别(ASR)评估,这是开放ASR排行榜首次加入全球南方语言。该项目包括针对印度英语的Monsoon en-IN公共(5.62小时)和私人(5.58小时)数据集,以及针对印地语的Monsoon hi-IN公共(1.33小时)和私人(4.47小时)数据集。这些数据集包含对话式、自发性语音,通过开放式叙述提示和后续问题在旅行、医疗、农业、教育和数字服务等领域进行引导,以促进更长的描述性交流。
Voice Arena and Hugging Face partner to launch open ASR evaluation for Hindi and Indian English
Benchmarks decide what gets built. A model that scores well on the Open ASR Leaderboard gets adopted and iterated on, while capabilities the leaderboard does not measure tend not to improve. Much of the recent work on the leaderboard has gone into making the evaluation metrics more trustworthy:
- Held-out private splits .
- Benchmark-fitting analysis to quantify how much models are reproducing reference transcripts rather than transcribing solely on the audio.
- Closing the gaps in normalisers to ensure correct predictions/variants are not penalized.
All of that makes one number (WER) harder to game. It is still one number. A long line of work has shown that ASR error rates are not evenly distributed across the people using them. Racial disparities in automated speech recognition found commercial systems roughly twice as bad for Black speakers as for white speakers, and Quantifying Bias in Automatic Speech Recognition found further differences by gender, age and accent. None of that is visible on a leaderboard, and not because the leaderboard is hiding it. The test sets it runs on record what was said and almost nothing about who said it.
To address this gap, we introduce two evaluation sets to the Open ASR Leaderboard: Monsoon en-IN and Monsoon hi-IN . Hindi, spoken by more than half a billion people, is the first Indic language on a multilingual tab that currently covers only European languages. Each set is released as a public split, available for self-scoring, and a private split withheld to limit benchmark-specific optimisation. The four splits are speaker-disjoint, comprising 4,888 speakers, with 12 speaker attributes recorded for each.
Design of the collection
A test set can only expose a failure mode it varies along. Most benchmarks are built from whatever audio was readily available. Monsoon was built to vary along nine axes: geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and the existence of multiple valid transcripts for the same audio. Each is a way an aggregate WER can be right on average and wrong for a particular population.
The collection method follows from that.
- Geography comes from recruiting across hundreds of districts rather than recording longer sessions in fewer places.
- Devices and acoustic conditions come from contributors using their own handsets and connections, indoors and out, rather than supplied hardware in a quiet room.
- Vocabulary, speech type and speech rate come from the prompts: everyday topics that push contributors toward opinion, disagreement, narration and recall, which is where named entities, numbers and unrehearsed phrasing appear.
- Age and gender are recorded per speaker and verified.
- Multiple valid transcripts is a property of the reference rather than the audio, and it is the subject of a later section.
Dataset composition
Four splits, two languages, collected through one pipeline.
Set Language Duration Speakers Clip length (mean / median) M/F Districts States/UTs Devices Style Transcription
Monsoon en-IN public Indian English 5.62 h 1,444 9.6s / 10.4s 50/50 428 24/6 556 Conversational, spontaneous Normalised, disfluencies
Monsoon en-IN private Indian English 5.58 h 1,405 9.6s / 10.4s 45/55 420 24/6 560 Conversational, spontaneous Normalised, disfluencies
Monsoon hi-IN public Hindi 1.33 h 468 6.4s / 5.0s 54/46 202 11/3 315 Conversational, spontaneous Lattice (accepted orthographic variants)
Monsoon hi-IN private Hindi 4.47 h 1,571 6.6s / 5.3s 55/45 295 12/3 582 Conversational, spontaneous Lattice (accepted orthographic variants)
The data is sourced from unscripted dual-channel spontaneous conversations, with clips segmented from a single channel so that each clip carries one speaker. Along with the fields reported in the table, each clip also records occupation, education, marital status, income band, handset brand, current city and years in the current district.
Five clips from the public Indian English split, with the metadata each one carries: