跳到正文
HNHacker News·
暂不在当前实时榜单

Show HN: JevBench, a reproducible benchmark for typed decision models

AI 摘要

JevBench v1.3.0, a reproducible benchmark for typed decision models, includes an updated run for Bespoke Nimble 9B. After Bespoke Labs increased the serving prompt limit from 2,048 to 8,192 tokens, the hard-tier accuracy for Nimble 9B rose from 43.6% to 65.5%. Despite this, the overall score fell due to changes in cost and speed, partly attributed to network distance. The benchmark also references classifier.dev's fast tier, which is Jev, priced at $0.0033 per 1,000 decisions under its Pro plan.

为什么是这条

This update to JevBench v1.3.0 provides a rare look into how a model's score can fall despite increased accuracy, unlike typical benchmark updates that only highlight improvements.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

收录当时偏移:UTC+02026年9月22日 21:01 UTC

收录
2026年9月22日 21:01
来源类型
未分类
爆款判定
判定依据
热度约为该来源近期上榜条目中位水平的 2.0 倍
指标对比
103 vs 中位 50.5(20 条基线样本)
检出时间
09/23 09:01

本站未收录正文。

前往源站阅读 →
来源·Hacker News·benchmarkheaven.com