Skip to content
HNHacker News·
Not on the current live radar

Show HN: JevBench, a reproducible benchmark for typed decision models

AI summary

JevBench v1.3.0, a reproducible benchmark for typed decision models, includes an updated run for Bespoke Nimble 9B. After Bespoke Labs increased the serving prompt limit from 2,048 to 8,192 tokens, the hard-tier accuracy for Nimble 9B rose from 43.6% to 65.5%. Despite this, the overall score fell due to changes in cost and speed, partly attributed to network distance. The benchmark also references classifier.dev's fast tier, which is Jev, priced at $0.0033 per 1,000 decisions under its Pro plan.

Why this one

This update to JevBench v1.3.0 provides a rare look into how a model's score can fall despite increased accuracy, unlike typical benchmark updates that only highlight improvements.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

IngestedOffset at this time: UTC+0Sep 22, 2026, 21:01 UTC

Ingested
Sep 22, 2026, 21:01
Source type
Unclassified
Breakout verdict
Basis
Running about 2.0× the median of this source's recent listed items
Metric comparison
103 vs median 50.5 (20 baseline samples)
Detected
09/23, 09:01

Full text isn't available here.

Read at source →
Source·Hacker News·benchmarkheaven.com