Contents · 1 / 5
01Abstract 02The null result 03Efficiency 04Frontier anchors 05Report and citation
JackBench · published study

First Pi–OpenCode Baseline: Kimi K3 Paid 4× More for the Same Quality

Abstract

This was the programme's first study, and it asked the basic question: do agent harnesses change outcomes for near-frontier models at all? Three models (GLM-5.2, Qwen 3.7 Max, Kimi K3) ran under Pi and OpenCode: 480 confirmatory attempts, 478 valid, over 16 task clusters, preregistered and deterministically graded. The pooled effect, OpenCode minus Pi, was +2.9 points with a 95% interval of −0.1 to +5.9 and p = 0.094: no general OpenCode advantage was established. The money told the story the pass rates hid: Kimi K3 tied exactly on quality whilst spending 4.06× more and using 19.4× the reasoning tokens under OpenCode. The takeaway: judge a harness on cost per success, because a quality tie can conceal a 4× difference in what you pay for it.

The null result

The study's primary contrast came back null, and it's reported as one. Per model, the OpenCode-minus-Pi effect was +1.25 points for GLM-5.2, +7.5 for Qwen 3.7 Max and exactly 0.0 for Kimi K3; pooled, the interval crosses zero. The heavier harness was not shown to buy quality for these models on this task set.

This baseline is what the rest of the programme is built against. The later Seven-Model Baseline found a real Pi advantage for one model under a Holm-corrected family; a contemporary replication of this study found its apparent advantages didn't survive cache-normalised safety scoring. Each of those is its own study with its own note. None of them are pooled with this one.

Efficiency

Quality parity hid a large operational difference, and it wasn't the same for every model.

Table 1 · OpenCode relative to Pi, per model
ModelΔ quality (pts)CostDurationReasoning tokensTotal tokens
Kimi K3 0.0 4.06× 1.84× 19.38× 3.56×
GLM-5.2 +1.25 1.95× 1.14× 0.98× 2.86×
Qwen 3.7 Max +7.5 1.11× 0.87× 0.70× 1.75×

Ratios are OpenCode over Pi. K3's row is the study's headline: the same quality at 4.06× the cost, with a 19-fold blowout in reasoning tokens.

Download the results (CSV)

Frontier anchors

Alongside the confirmatory grid, 96 anchor attempts across 6 frontier cells all passed deterministic checks. The report labels them exactly what they are: ceiling-limited descriptive references, not evidence of parity. They exist so later studies have an identity-proven frontier layer to calibrate against, reported separately from any model-only claim.

Report and citation

This study reached the programme's Published tier on 2026-08-07: a formal report and a locked result summary exist as a frozen bundle, and this page summarises them without re-deriving anything.

Tyler, J. (2026). First Pi–OpenCode Baseline (JackBench V3 V18). Cold Anvil Studios. https://coldanvil.com/research/first-pi-opencode-baseline/

@misc{tyler2026firstbaseline,
  author       = {Tyler, Jack},
  title        = {First Pi--OpenCode Baseline},
  year         = {2026},
  publisher    = {Cold Anvil Studios},
  howpublished = {\url{https://coldanvil.com/research/first-pi-opencode-baseline/}}
}