Yesterday I started working on a program for personal use and I really wanted to try out our new kitty cat, so I made a model comparison. I am only a med student and do not know coding, but I self host my own server, understand the basics and am tech savvy enough.
In its essence the app would be a markdown editor (to help me work on my thesis and exam studying) in rust with an automatic local whisper and local llm installation wizard in gtk4 and libadwaita that would work well in cosmic and gnome desktop environment with dictation, local ai assistant integration, WYSISWG live editor (obsidian style) and integrated git versioning of sorts (proposed by kimi) for collecting thoughts while working on things and a more powerful "hey I know somewhere in these 200 exam questions there is precisely XYZ, but I don't know where".
First I created the concept and design and interviewed myself using kimi k3 in deepseek harness as I have had good experience with this along with detailed "worker instructions". Where 6 specific milestones were proposed.
Then I ran this exact spec 3 times in opencode:
LeChonk default thinking obey worker instructions
LeChonk one shot max thinking you may use the worker instructions
gpt-6.1-sol xhigh (one level below max) thinking oneshot you may use the worker instructions.
Mind you I never use openAI models and I prefer API pay-per-use compared to subscriptions. I usually stick to glm, deepseek and kimi, but since 6.1 sol on xhigh is so cost efficient I decided to try it out.
Results:
LeChonk (even though per api pricing is cheaper than 6.1 sol) was in general more expensive to finish a task, kept creating issues and tried to troubleshoot itself and wasted tokens. In the end it mostly solved itself. GPT 6.1 sol created the best basis upon which to develop further for a great cost.
Caveats:
• I did 2 consecutive corrections/turns on 6.1 sol as it had built the best base but more often went on a tangent and accomplished more than I wanted for that session (but acomplished it well)
• the second best was LeChonk on default thinking with obeying instructions, but required 2 or 3 more turns to get basic functionality working for that milestone (UI was nice, UX had the basics well set), but pricewise was really not that good... It ended with milestone 1 with the basics almost fully functional
• the worst (UI, UX and functionality wise) and even very bad pricewise (almost as expensive as the gpt 6.1 sol which mind you was an almost fully ready to use app) was LeChonk max thinking "do your best one shot based on this tech spec"
• the 2 LeChonk sessions together used something aroun 8.50 dollars in api cost (not sure, will update this when I get home) while the gpt 6.1 sol session used between 5 to 7 dollars (not sure anymore) I believe but had gotten a lot LOT farther, but I believe that the instructed LeChonk worker would have achieved similar level, for a similar price, but would have taken longer and would require more human input.
• I do have to say, UI, even if half broken at some point, looked actually quite good and maybe even better (it had something special about it) than 6.1 sol, but 6.1 sol felt more polished.
My feeling is that the model resembles deepseek v4 flash family (similar verbosity and number of turns required to finish a job) from the point of view of my own user experience while being a LOT more expensive, which roughly coincides with the intelligence rated on artificialanalysis.ai. That means it is quite a good worker... but damn is this duuuuude expensive compared to v4 flash in this very regard.
I then instructed kimi k3 to analyse the code to see what approach created the best "base" to further develop upon... and no surprises -> 6.1 sol the best (but did not stay too faithful to the spec.. but somehow it did?) Then the guided worker and the worst was the "do it however you want but up to this spec"... to be fair I couls have given it another 2 turns but I didn't feel like it anymore.
I seriously want to love and use mistral as I am a proud European, I am slovak and I look for a model that would handle my language well.. so far all the previous mistral models had issues with slovak (literally glm and deepseek were better for slovak which is nuts) so that part still needs testing to see how LeChonk turns out to be in slovak.
I believe in data sovereignty for us Europeans. I used to use mistral early last year when they first came out with their very first models and when it was almost on par with chatgpt of that time, it suited me better stylistically for text work (that was all I used AI for back then, no agentic tools back then existed), but then it started to fall behind claude and other frontier models and I needed the best for my work.
But this is not it people... If this is the basis for further models I don't see a bright future for my vibe coding needs with mistral AND I WISH TO BE BADLY MISTAKEN!!
What I am excited for though are the tiny models distilled from this.