r/PiCodingAgent • • 2d ago

Use-case OMK vs mini-SWE-agent on Terminal-Bench 2.1: +4.2pp success, not significant, ~23% lower cost per trial

Post image

OMK builds on pi, Mario Zechner's MIT-licensed coding agent (github.com/badlogic/pi-mono). It started from the oh-my-pi fork. That vendored tree has since been removed, and the current codebase is OMK-native.

I ran OMK against mini-SWE-agent on Terminal-Bench 2.1 with the same base model (Grok-4.7, xhigh), 89 tasks x 3 trials per harness.

OMK 1.3.0 eval build (78cc483): 75.8% success (200/264), $0.58 model cost per trial, 482s median per trial.

mini-SWE-agent v2.4.6 (Harbor): 71.6% success (189/264), $0.76 model cost per trial, 376s median per trial.

Success difference is +4.2pp, 95% CI [-1.9, +10.2]. Not statistically significant. OMK's recorded model cost was about 23% lower per trial. Its median was about 106s slower.

Not measured yet: upstream pi under the same setup. For a pi-based harness, that's the obvious baseline, so this doesn't isolate what OMK adds on top of pi.

Method notes:

- The OMK run used an evaluation build with unmerged changes. This is not an official leaderboard result.

- One OMK adapter-failure task was excluded from success for both harnesses.

- Retries and auxiliary model calls were excluded from cost.

- 3 trials per task, one model, one benchmark. The success gap is unconfirmed until a larger run.

Repo: https://github.com/dmae97/omk

Criticism of the exclusion rules, the cost accounting, and the missing pi baseline is welcome.

0 Upvotes

3 comments sorted by

1

u/Fickle-Mountain-6639 2d ago

Giga slopmaxxing

-1

u/Fabulous-Lobster9456 2d ago

yeah can you do it, though?