WHAT-THE-JEV / RESOURCES
Decision models and the ecosystem taking shape around them.
A head-to-head benchmark for choice, noul, and score classification models, with two hash-locked suites.
A leaderboard comparing Jev and open decision models across more than 100 benchmarks.
A 14-dataset study of Jev as a reranker against Cohere, ZeroEntropy, and open models.
A blind security benchmark testing Jev on prompt injection and vulnerable-code detection.
Independent leaderboard scoring Jev and six LLMs on identical typed decision questions, with public per-decision probability data.
Solo independent benchmark comparing Jev with two GPT-6 models on two 500-sample datasets.
Benchmark Heaven's leaderboard for Jev-class decision models; four-axis geometric-mean score, cite with version.
Benchmark for typed decision engines with a full reproduction kit and a public-private mixed suite.
TypeSafe's official workflow-evals site; figures are vendor-reported, workflow code is open-sourced.