Keyboard shortcuts

⌘ / Ctrl K
Search documents
D
Toggle theme
L
Switch language
H
Go home
M
Copy article Markdown
Esc
Close dialog

WHAT-THE-JEV / RESOURCES

Benchmarks

Decision models and the ecosystem taking shape around them.

Benchmarks

01

classifier-benchmark (jabr)

A head-to-head benchmark for choice, noul, and score classification models, with two hash-locked suites.

↗
02

System One Mosaic Benchmark (S1MB)

A leaderboard comparing Jev and open decision models across more than 100 benchmarks.

↗
03

jev-rerank-bench (anessbelbati)

A 14-dataset study of Jev as a reranker against Cohere, ZeroEntropy, and open models.

↗
04

jev-sec-bench (Gaurav-Gosain)

A blind security benchmark testing Jev on prompt injection and vulnerable-code detection.

↗
05

Jevals.com

Independent leaderboard scoring Jev and six LLMs on identical typed decision questions, with public per-decision probability data.

↗
06

jev-bench (PavelRavich)

Solo independent benchmark comparing Jev with two GPT-6 models on two 500-sample datasets.

↗
07

JevBench (Benchmark Heaven)

Benchmark Heaven's leaderboard for Jev-class decision models; four-axis geometric-mean score, cite with version.

↗
08

Decision Index

Benchmark for typed decision engines with a full reproduction kit and a public-private mixed suite.

↗
09

evals.typesafe.ai

TypeSafe's official workflow-evals site; figures are vendor-reported, workflow code is open-sourced.

↗