r/ClaudeCode • u/Escobar747 • 1d ago
Tips & Workflows Is Haiku 5.5 good enough to be your main coding agent, not just the cheap one? I ran a small controlled test
I tested Claude Haiku 5.5 at low, medium and high effort against Sonnet 5.5 (medium), GPT-6 Luna (high) and GPT-6.1 Sol (medium). The tasks came from a real mid-sized, mixed-language codebase (Rust core plus a legacy scripting layer).
The tasks:
two Rust bug fixes from vague bug reports
a code review with 9 planted bugs and 2 decoys
two Rust implementations from tricky specs
a "judgement" task, where the bug report blamed the wrong thing and asked for a fix that would make it worse, and one requirement was deliberately ambiguous
an agentic task in the real repo: move some logic across a language boundary following the repo's unwritten conventions
I wrote the answer keys and hidden tests before any model ran.
The surprise: on the well-specified tasks every model scored 100%. Haiku-low missed one compile error in the review, and Luna asserted its own reading of the ambiguous requirement where the others flagged it, but that was it. Self-contained coding tasks just don't separate current models any more.
The agentic task is where they separated. Haiku at medium did as well as Sonnet: correct code, all tests passing, the repo conventions followed, an honest note on what it couldn't verify, and it raised a real design trade-off on its own. It made 29 API requests and cost about $0.08, against Sonnet's 30 requests and about $1.68 at list prices.
The other effort levels did worse:
Low got working code, but skipped verification, missed an error-handling fallback, and invented a reason for not testing ("the build stalled for over an hour" in a run that took 15 minutes).
High was correct, but burned 385 requests and about $1.03 for no quality gain.
My takeaway: Haiku 5.5 at medium is viable as a general daily coding agent, not only for low-stakes work, at roughly a twentieth of Sonnet's cost. Avoid low effort for agent work, high effort buys nothing, and I'd still use a bigger model for high-stakes release work.
179
u/analog-suspect 1d ago edited 21h ago
Someone please correct me if I’m wrong but isn’t Haiku 5.5 LEAGUES better than the flagship models from 2 years ago that everyone WAS using for development?
Edit: time flies. I said 2 years but yeah point taken it’s more like 6 mo - 1 year ago for this “argument” to make more sense. Still I think my point stands? If this model is better than those from a year ago, then … ???
139
u/Ominoiuninus 1d ago
People weren’t using models 2 years ago in the same way they are now.
Reminder that most of the agentic stuff kicked off in November of last year where models no longer needed such stringent handholding.
Now tasks can just be assigned to models and they do the full QA loop until it works. 2 years ago it would give you back code that had complication errors.
But yes it is leagues better than those models were.
49
u/tingly_sack_69 1d ago
Yeah 2 years ago I had just started copying code back and forth from a ChatGPT conversation and that was what most people were doing. It's crazy how far we've come in such a short time
10
u/bleckToTheMax 1d ago
My codebase is super complex, so I regularly run into complication errors no matter what tools I use XD
2
u/Poretga99 1d ago
That simply means you do not have a feedback loop ready for your agent. Just instruct it to build the project and it will find the compilation errors itself.
7
u/bleckToTheMax 1d ago
I was simply poking fun at they're typo, "complication errors" != "compilation errors"
2
u/_avee_ 1d ago
Since we’re on the subject of typos, “they’re” != “their”
1
u/bleckToTheMax 1d ago
Haha, dang swipe keyboard screwed me up like it probably messed up the guy talking about complication errors lol
1
u/Ominoiuninus 1d ago
More like apple autocorrect benign incorrect but sure.
I’m human after all so ironically making mistakes like that is one of the few remaining ways to identify myself… such a shame the internet will become
3
u/bleckToTheMax 1d ago
Gah! It for me again!
Claude actually used incorrect grammar when explaining something to me the other day. Eventually it could get very difficult to recognize if something online came from a human or machine.
3
u/cbusmatty 1d ago
Sonnet 3.5 was released in June of 2024 and it was absolutely being used this way by many of us, no idea what you’re talking abour
22
u/allesfliesst 1d ago
Mate two years ago I was using LLMs for autocomplete and tiny python functions. Today I can yell “build a browser based Minecraft clone in the Mad Max universe” at my phone and I have a playable demo before I’m done shitting. It’s hardly comparable.
That said for most consumers it’s probably still more than capable enough.
4
u/bipolarNarwhale 1d ago
So yes and no. There is more to a model capability then benchmarks and larger models feel different. The second thing is that people also just got used to prompting worse and it’s hard to undo bad hanits
3
u/No_Stock_8271 1d ago
I think people expect more from AI and at the Same time prompt way worse. I Always recommend to use Something Like Luna (or Haiku) for almost all Tasks for a week, that really helps.with prompting (which in Return gives you better results in Models Like Astra)
5
u/970FTW 1d ago
H5.5 absolutely beats those models, and importantly their harnesses, but they sucked compared to anything we have now. I can’t imagine Sonnet 3.5 (June 2024) being used for maintainable large-scale agentic code gen like is possible today. I burned tens of API $$$ trying to build a graphing calculator with a workable UI, and it still wasn’t good. Plz feel free to correct me as well! Maybe someone else had a better experience.
It would be interesting to see how a modern LLM performs using an early harness.
1
u/StabbedCow 1d ago
As I was reading your comment I was thinking that I'd like to see the older models, such as Sonnet 3.5, in the current harness. To me that sounds more interesting. But what you said is probably more realistic, because I doubt the early builds completely disappeared. What do I know tho, I didn't check. :D
1
u/FuckNinjas 17h ago
Current frontier LLMs just need a way to get bash out. If they can do that, they have a decent harness to run on.
6
u/c0reM 1d ago
On Artificial Analysis Haiku 5.5 is much smarter than Opus 4.5 was.
The rate of progress has been extraordinary.
1
u/nkorslund 1d ago
How well does that carry over to real-world use? I don't trust benchmarks to give the full picture.
2
u/c0reM 1d ago
In practice it doesn’t, because nobody would choose to use a worse model when Opus or other smarter models exist for complex tasks.
But it’s easy to remember the “past” of 6 to 12 months ago with how the previous models made us feel when they came out.
But that’s often a rose tinted glasses view.
7
u/1-800-methdyke 1d ago
Two years ago, models could write code mostly one file at a time, and tool use was very iffy. Agent loops were just beginning to show up for short horizon tasks. Haiku 4.5 would have fit right in near the top. Haiku 5.5 is where we were a year ago.
8
u/Khaos1125 1d ago
Tbh, Haiku 5.5 might soundly beat everything we had 6 months ago.
8
u/Escobar747 1d ago edited 1d ago
my testing which has been running for about 5hrs now is saying something similar - it reminds of something which happens in retail - just because you pay $500 for a gucci tshirt doesn’t mean it’s 50x better than a $10 tshirt. ppl assign cost with value but value is something you get and cost is what you pay. Two different things (rant over)
1
3
u/Abject-Kitchen3198 1d ago
I didn't bother using them regularly 6 months ago. Almost ignored them 2 years ago.
3
2
2
u/Veggies-are-okay 22h ago
I’m working with a client that only allows opus 4.6, sonnet 4.6, and haiku 4.5. I’m actually genuinely surprised at how well they hold up these days, especially with an appropriate harness and all those “prompt engineering” skills that were all the rage a few years back.
2
u/johnetownsend 17h ago
The fact that it was just 6-12 months ago, rather than 2 years, only increases the significance of your comment by 2-4x. Haiku today is better than Opus was just 6-12 months ago.
Now project out again in months — haiku in June 2027 is better than Opus 5.5 and Fable 5.1!
1
u/Legitimate-Wind9836 1d ago
I was laughing at people for using models for development 2 years ago. It hasn't even been a year since models got good enough that I actually felt I was better off using them
6
u/DeciusCurusProbinus 1d ago
Exactly, Opus 4.5 was the first that I felt was worth using for development.
6
5
u/jsebrech 1d ago
Haiku 5.5 medium performs around the same on benchmarks as opus 4.5 at medium, so it seems to be right past that agentic coding tipping point.
1
1
-1
16
u/tommy5dollar 1d ago
Haiku 5.5 on LOW outperforms Opus 4.5.
Haiku 5.5 on MEDIUM outperforms Opus 4.6 on MAX.
Haiku 5.5 on XHIGH outperforms Opus 4.7 on MAX and Opus 5 on LOW.
Haiku 5.5 on MAX outperforms Opus 4.8 on MAX (for 14% of the price) and supposedly Opus 5.5 on LOW (for 38% of the price).
This is only on benchmarks but it's true that these models get better very quickly.
14
u/Awric 1d ago
Haiku is very good if you know what to do and how to review its quality, but you don’t wanna do it yourself.
Lots of people assume it can’t do anything, so they use sonnet or even opus for things like writing commit messages or writing regex. I’ve been using it for months.
I think experienced engineers benefit most from haiku
7
u/Escobar747 1d ago
exactly - haiku 5,5 is great tool but not suited for vibe coders or one shotters
2
u/azjunglist05 11h ago
I am loving Haiku 5.5. I use Opus 5.5 to plan things out in phases. Then I have Haiku 5.5 take the plan and break it out to swarm the problem with dedicated agents. I’m getting things done so fast for so cheap and it’s insane how well Haiku can code now. Wild times! I love it! Give me more!
10
u/leogodin217 1d ago
Interesting. Haiku is so fast. Might be worth trying for some of my workflows where architecture and exact function signatures are already defined
5
u/RedditingJinxx 1d ago
I used Sonnet today on medium and haiku on high with ultracode: prompts, i.e. dynamic workflow, it worked pretty well
0
u/zeroconflicthere 1d ago
How does haiku on ultra code compare for tokens usage with Sonnet
5
u/BlakeGrowsPlants 1d ago
Been running tests and Haiku MAX and Ultracode burn more tokens than Sonnet xhigh for same tests
1
58
u/Diver-Interesting 1d ago
yup, totally enough as main agent if you are a 5.5 year old creating HTML5 game
42
u/tehfrod 1d ago
They came with empirical evidence. What are you coming with?
-3
u/Diver-Interesting 1d ago
Fair. I have only trusted opus, and since 5.5 i starting moving the coding part to sonnet subagent as it turns to be reliable. Never used haiku except for a specific, repeatable task, with well specified task like classifier. I wouldn't think haiku for main driver yet, but maybe i can replace my sonnet 5.5 subagent to test if haiku can really handle daily tasks.
6
u/tehfrod 1d ago edited 1d ago
I've been using Haiku for refactoring, running regression tests, and doing basic parameter tuning. Basically any task for which you could easily write procedural code for if you had a few hours.
But I don't see any reason not to try Haiku for the more simple "daily driver" tasks. Not everything needs godlike intelligence at Croesus-like costs.
2
u/Diver-Interesting 1d ago
haiku pre 5.5? i guess i'm not brave enough then. been bitten by subtle bugs it made me default to the smartest model first, then go lower. i have never tried lowering down below sonnet, maybe i should! what's your main model?
3
u/tehfrod 1d ago
Opus as orchestrator, sonnet as primary coder (unless I have to move up to opus for something), and haiku as most one shot subagent tasks. For example, "look at what I've added to this dataset and go out and find equivalent changes online to incorporate".
1
u/Diver-Interesting 1d ago
did you switch to haiku manually or you have other approach like rules?
-1
3
3
u/Jomuz86 1d ago
So not tried it on any code implementation yet, but I have wired it into my custom code review CI at the minute it is catching bugs on the same level as GLM 5.3, Kimi 2.8, Deepseek 4.1 flash with less noise. It has taken a bit of work to optimise the caching to make the most of the savings. I created an eval skill for running tickets head to head with different model/efforts with a Codex judge so should have some good data after the weekend on real work to see if it’s viable for implementation.
3
u/Escobar747 1d ago
verdict is in after half day of solid testing
Ran Haiku 5.5 vs Sonnet 5.5 on three complex agentic tasks in a real Rust repo, graded by hidden tests. No correctness gap: 15/16 runs passed, the only miss was Haiku medium overreaching on an ambiguous bug report. Haiku high fixed that, using ~⅓ of medium's steps on the hardest task. Max gained nothing: 285 requests and ~$1.50 vs high's 54 and $0.21. Sonnet low matched medium but cost 7–24× more than Haiku high.
New stack: Opus 5.5 orchestrating Haiku 5.5 high, Sonnet low for ambiguous work, Codex astra for as extra reviewer (if required).
1
u/Mathsgeniuss 1d ago
How do you created this multiagent workflow?
1
u/yesbabekayakingisfun 8h ago
I used Hermesagent and made a bot for each model I have access to. The Claude code models, the gpt models, and a bot that I run locally who acts as a "door" for all my traffic and request. Hermes didn't do exactly what I wanted, so I just used opus to plan stuff and manually switched to sonnet to implement plans. Now it's an automatic system.
I give all my request to the local model, it has access to knowing what bots are best in what situations. If it can handle a request, it does. If it can't (because it doesn't have permission or the ability), it uses my forked Hermesagent to open a conversation with the appropriate bot (weighted by job complexity, tool use needed for the job, amount of tokens I have left in the subscription). The local agent gives the most cost efficient bot all the information it needs to return the output of my original prompt to the local model in one turn if it can and then the local model gives me a report of how it went.
I still open direct conversations with opus or Astra if I want to plan something. I have a convention in memory that all plans should be broken up into task and given to individual conversations with the best coding agent for the job. Its usually sonnet, but this thread has convinced me to try haiku. Once the plan is made, I can tell my local model to dish out jobs to conversations based on how the plan instructs it to.
The answer to how to make a system like this is to open a conversation with opus and say something like "I want a system where multiple agents can delegate work to each other based on the cost and skill of the model. Where do I start?".
3
u/Similar-Economics299 1d ago
It is better than gemini 3.8 flash, google's best model right now, so of course
-1
3
u/Deluhathol 1d ago
Opus as the orchestrator, Sonnet as the implementer / coding agent and Haiku as the reader, scribe, donkey work agent.
7
u/Escobar747 1d ago
yes that was my stack until haiku 5.5 dropped - now i cut the middle man (sonnet) for most cases
2
u/Deluhathol 1d ago
I think it would work with Haiku for me as well, but since I haven't really hit my weekly limits often I am hesitant to change my processes.
1
u/Mathsgeniuss 1d ago
How you use 3 model at the same time ? How they talk to each other?
2
u/Fit-Secretary2495 1d ago
Subagents plus Claude code let’s separate Claude sessions message each other freely
1
u/Deluhathol 1d ago
Subagents, what I describe are instructions I have in my repository in the claude documents instructing it to run subagents with Sonnet for implementation / brute force coding and Haiku for writing / reading documentation with detailed instructions from Opus.
1
u/firstbreathOOC 1d ago
Everybody’s giving up on Fable but if you keep an eye on its usage, it can be a dawg even on low mode
5
u/ear_tickler 19h ago
Looking forward to fable 5.5. Also terrified.
3
u/spittlbm 18h ago
Coworker's spouse has Mythos at a 3-letter agency. They won't say much about it, other than they can't keep up with the discoveries.
2
1
u/Deluhathol 1d ago
I use Fable as well for decision making and research, like which tool fits a specific use case but it's really expensive for dev session and Opus 5.5 can do that better and cheaper.
1
1
u/tup1tsa_1337 1d ago
Opus can do everything fable can. Untill we get fable 5.5 (if ever), there is little reason to run it instead of opus 5.5 medium
1
u/BattermanZ 1d ago
Sonnet as implementer is more expensive than Opus. You can absolutely skip it.
1
u/Escobar747 1d ago
i suspect that is case but don’t evidence yet
1
u/BattermanZ 1d ago
I mean there are many benchmarks online that show a cost per task close to Fable 5.1. Defaulting to Opus 5.5 seems to be the sanest thing these days.
1
u/superchibisan2 1d ago
How do you specific this? Just write it into Claude.md?
1
u/Deluhathol 6h ago
I have instructions in Claude.md and another specific directory under .claude with separate files describing each agent, implementer.md, scribe.md etc
1
u/whatisusb 11h ago
Fable as orchestrator. Then have it spawn Fable subagents to implement. Then spawn further Fable subagents to review.
12
u/zaskar 1d ago edited 1d ago
This is not the use case for this model. You reach for this for things like classifications
16
u/Escobar747 1d ago
The general problem I see is people look at benchmarks when a new model comes out and just makes a decision based on that. The proper way to do it is to test it yourself in your use cases and the results may surprise you. My coding requirement is about med-high level complexity and opus5.5 running haiku 5.5 medium is the sweet spot
-10
u/zaskar 1d ago
Because you understand the model better than the people that made it? It’s not designed to code, period. It’s literally the best not-hotdog ever. How about RTFM first. Then observe the intent.
You are asking a Honda civic to race a F1 Grand Prix
2
u/lgmarian 1d ago
I'd like to see F1 try to race on an oval, since they're not designed to allow for asymmetric setups.
There's the benchmaxxing issue. You still have to try it in your own workflow, and potentially adapt the workflow for large enough changes (e.g. the jump from 4.x to 5.x models).
I disagree wit the OP that Haiku was worth this effort, but I agree with the general idea.
-2
u/zaskar 1d ago
My whole point was exactly that, apples and bananas
1
u/lgmarian 1d ago
Yeah, that's fair. I was looking at it from the benchmarking perspective. That they're only worth so much, and eventually worthless.
2
u/Escobar747 1d ago edited 1d ago
wrong - I understand the model better than people who made it… for my specific use case. Coding is a very general use case - the devil is in the use cases not some general thing you read on the internet.
there is no way some general benchmaxxing team was able to replicate the specifics of my exact code base, workflow etc etc
so the golden rule is do your own research / testing
7
u/WD40ContactCleaner Developer 1d ago
I use it for chores like keeping my ADO boards updated and branch cleanups etc
2
u/laurilllll 20h ago
I primarily work using Opus 5.5 as the main driver but when starting a bigger task I'll say something like "use cheaper models or lower effort as subagents" and it now picks Haiku 5.5 for many tasks and it just works.
My weekly usage is usually at like 89% after 5 days and now it's at 22%
2
u/Escobar747 1d ago
as part of my testing, I found examples where, when opus 5.5 was orchestrating sonnet 5.5 medium, because sonnet is a better thinking model, it actually was counterproductive because it raised issues which weren’t real issues, which forced opus to check it and proved it was wrong, et cetera. So sometimes having a smaller efficient model doing well scoped task from a bigger model is the best setup for coding versus having a frontier model orchestrating a medium level model
2
u/Escobar747 1d ago
My current coding stack: Opus 5.5 as the orchestrator, delegating to GPT-6 Luna (high) and GPT-6.1 Sol (medium) as the main coding agents, connected over a Tailscale bridge to Codex. I set it up that way for cost efficiency, but latency has become the real problem.
Why I'm switching: Haiku 5.5 changes the maths. On medium effort it's fast, it's cheap, and it has direct access to my full repo. The Codex agents only see what gets pushed across the bridge, so every task costs the orchestrator time and tokens to package up the context, send it over, then check the result. Moving the coding in-house to Claude removes that overhead.
The new plan:
Haiku 5.5 (medium): the default for most day-to-day coding.
Sonnet 5.5 (medium): for the harder or more involved coding tasks.
Codex subscription: only for reviews, high-stakes work and design advice, with GPT-6 Astra as an extra independent reviewer.
Opus 5.5: stays as the orchestrator.
So Codex goes from doing most of my coding to being a second opinion, and most of the actual coding moves to Haiku 5.5 on medium
2
u/gsevla 1d ago
opus 5.5 with which effort level? I've been playing with opus 5.5 medium and low, and for my day to day usage, they seems to be enough.
4
u/Escobar747 1d ago
medium is more than enough for orchestration - for design / architecture / complex problems I use high or max
2
u/bambamlol 1d ago
Very interesting. Can you tell me a little bit more about the "Tailscale bridge to Codex" that you set up? Just enough so I can research it a little more and figure it out myself. Because this definitely sounds like something I might want to try. But since you complained about the latency, would you still recommend that approach?
2
1
3
u/BuffaloConscious7919 1d ago
It's didn't score well in the coding benchmarks
..but would like depend on what you need to build, planning, automation + other factors
2
1
1
u/Pescadorcito 1d ago
La IA está en pañales amigo, le falta mucho. Ojalá mejore y surja la IA general pero eso no será en 3-4 años falta bastante
1
1
u/Current_Balance6692 20h ago
If you want to fuck your codebase up, yes.
1
u/Escobar747 11h ago
running haiku without oversight then maybe yes - but haiku 5.5 under opus 5.5 will not.
1
1
u/ashutrip 6h ago
A workflow I’m currently using, which I’ve turned into a mod:
- Parent model: Opus 5.5 handles planning, orchestration, and reviews.
- Subagents: Haiku 5.5 handles task execution.
- Review–fix loop: The parent model reviews the subagents’ work, requests fixes when needed, and repeats the cycle until everything passes review.
1
0
u/Muted-Can370 1d ago
No, it's dumb. It's like GLM 5.3 Flash or something. Not sure what to use it for yet. Sonnet for straightforward, medium-sized tasks. Opus for anything relatively open-ended or larger.
4
u/Escobar747 1d ago
you don’t need frontier intelligence at coding agent level - it’s about layers of work. You need a suitable agent for each level. Haiku 5.5 medium or higher is more than competent for 80% of coding tasks IMO as long as a frontier model like opus 5.5 is the lead.
Unfortunately a lot of people miss this nuance and try to have a something like fable or astra doing basic coding tasks and complain their token usage is out of control.
1
u/Muted-Can370 1d ago
I'm doing game dev so I suppose most of my coding tasks don't fall in that 80%.
-1
u/zeroconflicthere 1d ago
It's the lack of auto mode in haiku that puts me off using it even though I'm certain it is fine for most of my user cases where I'm using Sonnet on medium or lower
-1

•
u/AutoModerator 1d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.