Positioning Jevals.com is an independent third-party leaderboard that runs TypeSafe's Jev decision model and six large language models through the same set of decision questions, producing directly comparable results for Jev and mainstream LLMs side by side.

What it does Every question uses the typed Noul, Choice, and Score formats that Jev answers natively, and outputs are scored against human labels from the PubMedQA, Banking77, and HelpSteer2 datasets. Six large language models answer the same questions alongside Jev. The site also publishes the raw per-decision probability data behind the leaderboard under a CC BY 4.0 license, so readers can recompute aggregate scores or run their own analyses instead of taking the headline numbers on faith.

Characteristics The leaderboard was still expanding at publication time, so task coverage and rankings should be expected to change. Its first published conclusion (2026-09-18, independent source) reported that on Noul questions Jev was statistically tied with the strongest LLM tested, at roughly 1/28 of the price. It is one of the few public sources of comparable per-decision probability data covering both Jev and mainstream LLMs, and any figure should be read as tied to the site's current version and update date rather than as a fixed result.

When to use Suited to readers who need same-question comparison data between Jev and mainstream LLMs, or per-decision probability data for secondary analysis, and to anyone who wants a reference source when designing typed-question evaluations of their own. Before citing a figure, check the site version and update date it corresponds to, since coverage and rankings may have shifted since the figure was first published.