Keyboard shortcuts

⌘ / Ctrl K
Search documents
D
Toggle theme
L
Switch language
H
Go home
M
Copy article Markdown
Esc
Close dialog

WHAT-THE-JEV / EXPERIMENTS

Experiments

Benchmarks and investigations, with data, results, and analysis.

1. LLM Benchmarks: Academic Ability

01

GPQA Diamond Graduate-Level Science

Explore JEV's performance on GPQA Diamond graduate-level science questions.

198 samplesCost $0.004594548
↗
02

GSM8K Grade-School Math Word Problems

Explore JEV's performance on GSM8K grade-school math word problems.

1,319 samplesCost $0.025092774
↗
03

MMLU-Pro Multi-Discipline Questions

Explore JEV's performance on MMLU-Pro college-level questions across 14 disciplines.

12,032 samplesCost $0.279330744
↗

2. LLM Benchmarks: Social Commonsense and Bias

04

BBQ Bias-Sensitive QA

Explore JEV's performance on BBQ's bias-sensitive questions, and whether it leans toward stereotypes.

58,492 samplesCost $0.921154332
↗
05

SocialIQA Social Commonsense

Explore JEV's performance on SocialIQA social commonsense questions.

2,224 samplesCost $0.035755734
↗

3. Paper QA System

06

Paper Classification

Explore whether JEV can decide, based on a research preference, whether an arXiv paper belongs in a paper library, and which topic it falls under.

400 samplesCost $0.016174452
↗
07

Paper QA

Explore whether JEV can answer questions about a paper based on its own text.

240 samplesCost $0.066836868
↗
08

Prompt Routing

Explore whether JEV can decide from the question text alone whether a question in paper QA goes to the JEV fast path or the LLM slow path.

480 samplesCost $0.019004748
↗

4. Classical Machine Learning Prediction: Regression and Classification

09

Boston Housing: Prediction and Comparison

Explore whether JEV can predict Boston suburb housing prices, and compare price levels across suburbs.

1,106 samplesCost $0.043098972
↗
10

Titanic Survival

Explore whether JEV can predict whether a Titanic passenger survived from the passenger record, and compare structured fields with prose text.

1,782 samplesCost $0.040420086
↗

5. LLM Post-Training Annotation

11

DPO Preference Annotation

Explore JEV's performance on preference annotation for DPO training data.

300 samplesCost $0.0300741
↗
12

GRPO Reward Assignment

Explore JEV's performance on GRPO trajectory reward assignment.

300 samplesCost $0.085096116
↗

6. Skill Routing

13

Skill Routing

Explore whether JEV can route user tasks to the right skill, skill set, or none, among 126 real agent skills.

500 samplesCost $0.214509582
↗