HNHacker News·
暂不在当前实时榜单
Show HN: JevBench, a reproducible benchmark for typed decision models
JevBench v1.3.0, a reproducible benchmark for typed decision models, includes an updated run for Bespoke Nimble 9B. After Bespoke Labs increased the serving prompt limit from 2,048 to 8,192 tokens, the hard-tier accuracy for Nimble 9B rose from 43.6% to 65.5%. Despite this, the overall score fell due to changes in cost and speed, partly attributed to network distance. The benchmark also references classifier.dev's fast tier, which is Jev, priced at $0.0033 per 1,000 decisions under its Pro plan.
This update to JevBench v1.3.0 provides a rare look into how a model's score can fall despite increased accuracy, unlike typical benchmark updates that only highlight improvements.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
收录当时偏移:UTC+02026年9月22日 21:01 UTC
- 收录
- 2026年9月22日 21:01
- 来源类型
- 未分类
- 判定依据
- 热度约为该来源近期上榜条目中位水平的 2.0 倍
- 指标对比
- 103 vs 中位 50.5(20 条基线样本)
- 检出时间
- 09/23 09:01
本站未收录正文。
前往源站阅读 →