Show HN: JevBench, a reproducible benchmark for typed decision models

Coding & Dev · Free · collected from public sources

Visit Show HN: JevBench, a reproducible benchmark for typed decision models ↗

What is Show HN: JevBench, a reproducible benchmark for typed decision models?

Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects <i>really</i> perform in comparison.<p>Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on.<p>JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting.<p>A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost.<p>Leaderboard right now:<p><pre><code> #1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3. </code></pre> MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes:<p><a href="https:&#x2F;&#x2F;github.com&#x2F;fstandhartinger&#x2F;jevbench" rel="nofollow">https:&#x2F;

Tags

ai hackernews show-hn

Pricing

Listed as Free. Pricing changes often — confirm on the official site.

Similar Coding & Dev tools

Loading…

Listings are collected automatically from public sources and refreshed daily. We do not take payment for placement.

← Back to the directory