Show HN: JevBench, a reproducible benchmark for typed decision models
JevBench v1.3.0, a reproducible benchmark for typed decision models, includes an updated run for Bespoke Nimble 9B. After Bespoke Labs increased the serving prompt limit from 2,048 to 8,192 tokens, the hard-tier accuracy for Nimble 9B rose from 43.6% to 65.5%. Despite this, the overall score fell due to changes in cost and speed, partly attributed to network distance. The benchmark also references classifier.dev's fast tier, which is Jev, priced at $0.0033 per 1,000 decisions under its Pro plan.
This update to JevBench v1.3.0 provides a rare look into how a model's score can fall despite increased accuracy, unlike typical benchmark updates that only highlight improvements.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Sep 22, 2026, 21:01 UTC
- Ingested
- Sep 22, 2026, 21:01
- Source type
- Unclassified
- Basis
- Running about 2.0× the median of this source's recent listed items
- Metric comparison
- 103 vs median 50.5 (20 baseline samples)
- Detected
- 09/23, 09:01
Full text isn't available here.
Read at source →