LocalsOnlyevaluationsHow to read this
THE EXPERIMENT SHELF · 6 BENCHMARKS

Put a little AI
to the test.

Six tests. Real local runs. Find out what each model can do.

Find your experiment

Curiosity encouraged ↙

Measured in our lab

Coding14 results

Terminal-Bench

Can it finish a task in a real terminal?

MEASURED30.3%

GB10 · 128GB

Explore benchmark
Results · 14 recorded
What the research adds

The paper separates execution, coherence and verification failures. A task can fail because the environment or tools break, because the agent loses track of the goal, or because it fails to check its work.

Terminal-Bench research
Coding11 results

SWE-bench Pro

Can it fix an issue in an existing codebase?

MEASURED20.0%

GB10 · 128GB

Explore benchmark
Results · 11 recorded
What the research adds

The failure analysis distinguishes incorrect solutions, syntax errors, wrong files, instruction following, tool use, context limits and loops. The diagnosis is richer than a single repository-fix score.

SWE-bench Pro paper
Reasoning15 results

GPQA Diamond

Can it reason through difficult science?

MEASURED73.2%

GB10 · 128GB

Explore benchmark
Results · 15 recorded
What the research adds

Expert-written biology, chemistry and physics questions measure difficult scientific question answering. They do not establish practical experimental or research competence.

GPQA paper