Turning GLM-5.3-Flash into a Jev-like decision model
The study transforms GLM-5.3-Flash into a Jev-like decision model, comparing its performance against Jev and Laya across 28 datasets. While a Wilcoxon signed-rank test showed no significant difference (p = 0.64), the number of options impacted accuracy differently. On TREC, Jev dropped from 92.1% to 85.6% with increased options, GLM-5.3-Flash from 91.2% to 79.6%, and Laya from 88.4% to 51.2%. Conversely, on MASSIVE, Jev and GLM-5.3-Flash improved, while Laya declined, indicating task difficulty isn't solely determined by option count.
Unlike previous studies, this report specifically compares GLM-5.3-Flash against Jev and Laya, showing how option count impacts their performance differently across datasets like TREC and MASSIVE.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月27日 00:00 UTC
- 收录
- 2026年9月27日 00:00
- 来源类型
- 未分类
本站未收录正文。
前往源站阅读 →