Introducing Falcon ASR
Falcon-ASR is a 1.6 billion parameter speech recognition model developed at the Technology Innovation Institute (TII) in Abu Dhabi. It focuses on Arabic, particularly the Emirati dialect, achieving an Arabic WER of 20.92% and an Emirati WER (TII evaluation) of 22.73%. The model also supports English, French, Spanish, and Portuguese.
This is the first 1.6 billion parameter speech recognition model from TII, offering specific support for the Emirati dialect unlike previous models.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年10月7日 13:21 UTC
收录当时偏移:UTC+02026年10月8日 14:00 UTC
- 发布
- 2026年10月7日 13:21
- 收录
- 2026年10月8日 14:00
- 来源类型
- 官方发布
- 档位
- 当事方
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
Arabic WER: 20.92% · Parameters: 1.6B · Emirati WER (TII evaluation): 22.73%
We’re introducing Falcon-ASR, our 1.6 billion parameter speech recognition model for Arabic, with a particular focus on the Emirati dialect. Developed at the Technology Innovation Institute (TII) in Abu Dhabi, it also supports English, French, Spanish and Portuguese.
In our evaluation, Falcon-ASR achieved an average word error rate of 20.92% across six Arabic test sets, compared with the best published result of 23.17% in the leaderboard snapshot we used. On our internal Emirati evaluation, it recorded the lowest word and character error rates among the systems we compared.
We also support word-level timestamps for transcriptions, linking each transcribed word to its position in the audio.
Try Falcon ASR →
Recognising spoken Arabic
Arabic speech varies by region, speaker and setting. A model that handles a formal news broadcast may still struggle with a conversation in Emirati or with speech recorded over a phone line. Dialectal Arabic also has fewer transcribed resources than Modern Standard Arabic (MSA), which makes training and evaluation harder.
We trained Falcon-ASR on Emirati, MSA, other Gulf and Arabic dialects, and English. Our aim is to transcribe the words people use in everyday speech, including dialectal forms and changes between languages.
Arabic benchmark results
The Open Universal Arabic ASR Leaderboard, maintained by the ELM Research Center, ranks systems by the equal-weight average WER across six test sets. It also reports character error rate (CER). Lower values are better for both metrics. Our Falcon-ASR evaluation follows this protocol.
WER = Word Error Rate; CER = Character Error Rate. A lower value indicates better performance.
We evaluated Falcon-ASR on the same six benchmarks using the leaderboard’s pinned manifests. Competitor figures are the published leaderboard averages checked on 30 September 2026. Falcon-ASR’s average WER is 2.25 percentage points better than the best published result in that snapshot.
Evaluating Emirati speech
Public evaluation data already includes Emirati: Casablanca has a UAE subset. We complement that coverage with an internal evaluation of additional Emirati and Gulf speech, using held-out recordings and human-validated transcripts to assess transcription accuracy beyond the public UAE subset.
In our internal Emirati evaluation, Falcon-ASR achieved 22.73% WER and 10.19% CER:
Falcon-ASR has the lowest WER and CER among the systems compared here. Its WER is 4.07 percentage points below Qwen3-Omni, the next best result. The results show improved transcription accuracy at both the word and character level on this evaluation.
Training for different recording conditions
We included background noise, overlapping speech, music, room reverberation and telephony effects, as well as variations in speed and pitch. We applied the same treatment to Emirati recordings, exposing the model to a range of conditions it may encounter in meetings, calls and other everyday recordings.
English and other languages
Falcon-ASR also transcribes English with the same model weights. In our evaluation on the seven public English test sets used by the Hugging Face Open ASR Leaderboard, it achieved a mean WER of 5.74%.
The model also supports French, Spanish and Portuguese. All five languages use the same weights, without requiring a language flag. The output is a transcript in the language spoken.
Model foundation
Falcon-ASR builds on our Falcon3-Audio work. The architecture and training approach for Falcon3-Audio are described in Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data.
Try Falcon ASR