Contents · 1 / 7
Seven-Model Pi–OpenCode Baseline: Harness Effects Are Real and Model-Specific
Agent harnesses are usually chosen on reputation, and reputations assume the harness works the same for every model. This study asked whether that assumption holds. Seven hosted models ran the same 72 tasks in Pi and OpenCode: 1,008 planned attempts, 1,003 effective, a scored sample of 153 USD across 7,255 provider requests, deterministic grading throughout. One harness effect survived Holm correction: GPT-5.6 Sol gained 19.4 points in Pi (58/72 against 44/72, adjusted p = 0.009). The other six models showed no corrected effect in either direction. So the harness effect is real, and it belongs to the pairing rather than the harness. The takeaway: a harness that transforms one model can do nothing for yours, so measure your own model-harness pairing before trusting anyone else's ranking. This is a standalone study inside the JackBench programme; the local-model counterpart is the Local Qwen 3.6 Harness Baseline, and the two are never pooled.
The question
The registered research question: when model, task and feasible operating conditions are fixed, how much does harness choice change agent outcomes? A secondary question about the frontier–near-frontier gap shares this study's evidence; its frontier layer is reported separately and never as a model-only claim.
Method
Each model ran the same 72 tasks in each harness. Grading is deterministic, every attempt writes a sealed receipt, comparisons were registered in advance, and every reported p-value carries a Holm correction across the registered family. Of 1,008 planned attempts, 5 were lost to provider failures; the effective sample is 1,003.
1,034 physical receipts including replacements; 32 provider-invalid attempts (1.08 USD) are excluded, and reported by identity and reason.
The grid
Figure 1 as a table
| GPT-5.6 Sol | Pi 80.6% | OpenCode 61.1% | Holm p = 0.009 |
| Opus 5 | Pi 72.2% | OpenCode 73.6% | p = 1.0 |
| GLM 5.2 | Pi 65.3% | OpenCode 52.8% | p = 0.294 |
| Qwen 3.8 Max | Pi 61.1% | OpenCode 63.9% | p = 1.0 |
| K3 | Pi 59.7% | OpenCode 61.1% | p = 1.0 |
| Qwen 3.7 Max | Pi 54.2% | OpenCode 45.8% | p = 0.898 |
| Fable 5 | withheld, provider missingness | ||
Best cell in the grid: Sol on Pi at 80.6%. The model-specific spread is the finding: the same harness change helped Sol by 19 points and did nothing measurable for the other six.
Download the results (CSV)Efficiency
K3 succeeded at near-identical rates in both harnesses whilst spending 5.62× more and taking 3.22× the median duration in OpenCode. Judged on solve rate alone the two set-ups look interchangeable for K3, which is exactly why the study records cost and duration as well.
Two more efficiency observations from the sealed record: Qwen 3.8 Max and Opus 5 gained no confirmatory quality from their more expensive OpenCode cells, and Sol's Pi cell cost slightly more than its OpenCode cell whilst producing 14 additional successes and the only adjusted-significant advantage.
Corrections and limitations
Known limits: two harnesses, one task set, and provider missingness that withheld Fable 5's comparison. Hosted results don't license claims about harnesses this study never ran; those are being measured separately and are listed on the research page with their own statuses.
Cite this
Tyler, J. (2026). Seven-Model Pi–OpenCode Baseline. Cold Anvil Studios. https://coldanvil.com/research/seven-model-pi-opencode-baseline/
@misc{tyler2026sevenmodel,
author = {Tyler, Jack},
title = {Seven-Model Pi--OpenCode Baseline},
year = {2026},
publisher = {Cold Anvil Studios},
howpublished = {\url{https://coldanvil.com/research/seven-model-pi-opencode-baseline/}}
}