Coding & Dev · Free · collected from public sources
Visit Show HN: JevBench, a reproducible benchmark for typed decision models ↗
Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects <i>really</i> perform in comparison.<p>Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on.<p>JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting.<p>A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost.<p>Leaderboard right now:<p><pre><code> #1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3. </code></pre> MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes:<p><a href="https://github.com/fstandhartinger/jevbench" rel="nofollow">https:/
ai hackernews show-hn
Listed as Free. Pricing changes often — confirm on the official site.
Listings are collected automatically from public sources and refreshed daily. We do not take payment for placement.