LocalsOnlyevaluationsHow to read this
← Model shelfQwen / Alibaba

Qwen 3.8 27B

Setups

Local recipes

Unspecified references

Comparison details

These records do not establish a matched comparison protocol. A rank requires the same benchmark, scoring metric, protocol, sample, and evidence type within an execution group. Different tests are never averaged into a recipe rating. Hardware, precision, and evaluation settings may differ; comparisons do not isolate quantization effects.

ON THIS SETUP · MEASURED BY LOCALSONLY

UD-Q4_K_XL

GB10 · 128GB · llama.cpp

6 benchmark results
Coding

LiveCodeBench

Solve programming problems with executable tests.

Scoring & limits

Version 6.0 · 100 tasks

Generated solutions are checked by executable tests. The benchmark version and problem window matter; this is different from maintaining an existing repository.

Test pass rate66.0%66 / 100 passed
Measured · 2026-09-0411.5 t/s generationInspect evidence ↗
Coding

SWE-bench Pro

Resolve issues in real software repositories.

Scoring & limits

Pro evaluation · 50 tasks

Repository tests verify issue resolution. Task subsets should only be compared when their selection and evaluation protocol match.

Issues resolved20.0%10 / 50 passed
Measured · 2026-09-0310.8 t/s generationInspect evidence ↗
Coding

Terminal-Bench

Multi-step tasks in a terminal environment.

Scoring & limits

Version 2.1 · 89 tasks

The task must pass its verification in a terminal environment. Tools, harness, context budget and execution failures all contribute to the result.

Tasks completed30.3%27 / 89 passed
Measured · 2026-09-026.8 t/s generationInspect evidence ↗
Instructions

IFBench

Follow verifiable instruction constraints.

Scoring & limits

Loose prompt grading · 300 tasks

This receipt uses loose prompt-level grading. It is not the strict score, and the local and published protocols are not matched.

Loose prompt accuracy42.3%127 / 300 passed
Measured · 2026-09-08Inspect evidence ↗
Office & CRM

AutomationBench

Sales, marketing, operations, support, finance and HR workflows.

Scoring & limits

Local evaluation · 300 tasks

A workflow is complete only when every assertion passes. Mean reward gives partial credit for progress; it is not the share of jobs finished. Task sets and evaluation budgets can differ between receipts.

Mean reward 48.8% · partial credit, not completed jobs.

Fully completed18.3%55 / 300 completed
Measured · 2026-09-0510.1 t/s generationInspect evidence ↗
Reasoning

GPQA Diamond

Graduate-level science questions.

Scoring & limits

Diamond · version 1.0 · 198 tasks

Credit comes from answering the science question correctly. Multiple-choice accuracy does not establish research or experimental competence.

Answer accuracy73.2%145 / 198 passed
Measured · 2026-09-0410.2 t/s generationInspect evidence ↗
From scores to real attemptsAutomationBench · follow a pass and a miss

BEYOND THE HEADLINE

What happened across 300 tasks?

Each block represents one recorded task, in receipt order from left to right, then top to bottom. Fill shows the outcome; work-category colors are not used here.

Fully passed55 18.3%
Partial progress149 49.7%
No credit96 32.0%

Tested by LocalsOnly · UD-Q4_K_XL · GB10 · 128GB. Partial progress still requires further work before a task is complete.

Inspect individual task outcomes →
Completed

Update contact phone · sales

Completed the recorded checks.

4 / 4 assertions · 13 steps

Follow this exact task ↗
Exact task metadata

sales.update_contact_phone · ID 1

sales · 4 / 4 assertions · 13 steps / 24 tools · passed

Incomplete

Create important draft · sales

The attempt stopped before completion. The record does not establish why.

0 / 2 assertions · 26 steps

Follow this exact task ↗
Exact task metadata

sales.create_important_draft · ID 9

sales · 0 / 2 assertions · 26 steps / 98 tools · aborted_rollout

Recipe, sources & reproducibilityExact artifacts and evaluation provenance
Runtime
llama.cpp
Model host
ASUS GX10 · 128GB
Weight revision
4ca720788d1e01f1bff70c033e0d0028fd02e502

Context, sampling, harness settings and dataset versions can vary by benchmark. Each evidence link above contains its run-specific configuration.

Download this setup’s recipe JSON ↓