跳到正文
Bblog·
Archived topic · 归档话题,来源已停止追踪

QIMMA قِمّة ⛰: A Quality-First Arabic LLM Leaderboard

AI 摘要

QIMMA is a new Arabic LLM leaderboard that prioritizes quality validation of benchmarks. It addresses issues like fragmented evaluation, translation problems, and lack of quality checks in existing Arabic NLP evaluations. QIMMA systematically validates 109 subsets from 14 source benchmarks, covering 7 domains and over 52,000 samples, ensuring 99% native Arabic content and including the first Arabic leaderboard with code evaluation. Its multi-stage validation pipeline, involving LLMs and human review, revealed systematic quality problems in widely-used benchmarks, leading to the discarding of problematic samples.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年7月5日 04:00 UTC

收录
2026年7月5日 04:00
来源类型
未分类