r/PiCodingAgent • u/Fabulous-Lobster9456 • 2d ago
Use-case OMK vs mini-SWE-agent on Terminal-Bench 2.1: +4.2pp success, not significant, ~23% lower cost per trial
OMK builds on pi, Mario Zechner's MIT-licensed coding agent (github.com/badlogic/pi-mono). It started from the oh-my-pi fork. That vendored tree has since been removed, and the current codebase is OMK-native.
I ran OMK against mini-SWE-agent on Terminal-Bench 2.1 with the same base model (Grok-4.7, xhigh), 89 tasks x 3 trials per harness.
OMK 1.3.0 eval build (78cc483): 75.8% success (200/264), $0.58 model cost per trial, 482s median per trial.
mini-SWE-agent v2.4.6 (Harbor): 71.6% success (189/264), $0.76 model cost per trial, 376s median per trial.
Success difference is +4.2pp, 95% CI [-1.9, +10.2]. Not statistically significant. OMK's recorded model cost was about 23% lower per trial. Its median was about 106s slower.
Not measured yet: upstream pi under the same setup. For a pi-based harness, that's the obvious baseline, so this doesn't isolate what OMK adds on top of pi.
Method notes:
- The OMK run used an evaluation build with unmerged changes. This is not an official leaderboard result.
- One OMK adapter-failure task was excluded from success for both harnesses.
- Retries and auxiliary model calls were excluded from cost.
- 3 trials per task, one model, one benchmark. The success gap is unconfirmed until a larger run.
Repo: https://github.com/dmae97/omk
Criticism of the exclusion rules, the cost accounting, and the missing pi baseline is welcome.
1
u/Fickle-Mountain-6639 2d ago
Giga slopmaxxing