Published reference2026-02-19
SWE-bench Verified
swe-bench-verified · 500 tasks. Compare only matched task sets and protocols.
Gemini 3.1 Pro Preview · Google evaluation · Google DeepMind · High thinking · Google scaffold
↗ Published resultSWE-bench Verified
Gemini 3.1 Pro model card · SWE-Bench Verified ↗Tokens per second · see measurement basis
As recorded on this receipt
Run configuration & provenance +
- Context
- Not recorded
- Temperature
- Not recorded
- Wall time
- Not recorded
- Weight file
- Provider-managed / not reported
- Revision
- Not reported
- Reasoning
- Not reported
- Agent version
- Not reported
- Output token limit
- Not reported
- Dataset revision
- Not reported
- Harness host
- Not reported
- Model host
- Published reference
- Input tokens
- Not reported
- Output tokens
- Not reported
- Cached input tokens
- Not reported
- Agent turn / step limit
- Not reported
- Concurrency
- Not reported
- Weight file size
- Not reported
Throughput measurement: Measurement method not separately reported. This should not be treated as standardized decode-only speed.
Google scaffold · single attempt · bash + file-operation + submit tools · 10-run average. Thinking High. Score includes +0.6% for three official harness bugs Google documented. n=500.
Google model card: SWE-Bench Verified agentic coding, single attempt, 80.6% (Thinking High). Methodology: https://deepmind.google/models/evals-methodology/gemini-3-1-pro — 10x average; +0.6% adjustment after Gemini passed three items that official harness bugs made impossible. Not copied from a DeepSeek comparison table. Terminal-Bench on this Google recipe is 2.0, not 2.1. Not a LocalsOnly desk run.
Run ID: research-gemini-3.1-pro-preview-google-evaluation-swe-bench-verifiedTask explorer
Passed tasks first, then other outcomes grouped by recorded issue. Use the filters to inspect a particular outcome.
No task records in this receipt.
The headline or summary is available above. Individual outcomes have not been provided.
Explore test coverage →