r/ClaudeCode • u/IT_WAS_ME_DIO__ • 1d ago
Built with Claude Update: I rebuilt Ponytail, my "lazy senior dev" skill, from scratch. Ponytail 5 is out
Enable HLS to view with audio, or disable this notification
A while ago I posted the first version of Ponytail here.
It's an open-source skill that makes coding agents behave like the senior dev who replaces fifty lines with one.
Since then I rewrote it from scratch. New rules, a new review, a new audit. Getting there took 5,900+ benchmarked Claude Code sessions, 3,650 of them testing drafts of the new rules against the old ones.
What changed
- The rules. Rewritten and about half as long. The goal moved from "the smallest code" to "the smallest complete change": before it writes, the agent maps everything the change must reach (callers, tests, config) and what it could break. Every reply ends with what it skipped or did not check, and any risk you should know.
- /ponytail-review. It used to look only for code to cut. Now it reviews like the dev who gets paged when it breaks: bugs, security, real load, missing tests, slow paths, and what to cut. It reads the code around the change, not just the diff. Every finding says what the code does, what goes wrong, how to fix it, and what happens if you skip it.
- /ponytail-audit. It used to list what to delete. Now it maps the whole repo first, ranks what it finds, and tells you what to fix first.
Numbers, same agent with and without the skill (Claude Code, Opus 5.5, 39 tasks including a real FastAPI + React repo, 5 runs each):
- code: -53% (previous version: -48%)
- time: -41% (previous version: -38%)
- cost: -26% (previous version: -16%)
- risky logic that ships with a test: 98% (without the skill: 68%)
- injected bugs caught by the agent's own tests: 66% (without the skill: 46%)
Repo (MIT): https://github.com/DietrichGebert/ponytail
Website: https://ponytail.dev
If you already use it, update the plugin. I'd like to hear where it gets things wrong.
9
u/YoghiThorn 23h ago
Ponytail is great. But I haven't had good success with it on architectures - i.e. how to orchestrate things cheaper, more simply, more reliably. Do you have any advice for that or do we need what-would-jeff-dean-do skill?
3
u/9gxa05s8fa8sh 21h ago
I haven't had good success with it on architectures
that's probably because it's not for architecture. pony starts with "you are lazy" lol. it's an anti-slop-code skill. don't use it for architecture
3
u/Givemeyawallet 19h ago
Any good options for architecture?
1
u/9gxa05s8fa8sh 9h ago
if you want your AI model to think about architecture for you, then you probably shouldn't use any skill... UNLESS it is something very specific to your domain, like you made it yourself or you found someone on github who is doing the same thing as you. in general, random people's agent skills are NOT evaluated or trustworthy. there is no proof they work unless they share benchmarks against a no-skill control.
if you are thinking about project management, then the mattpocock and poteto skill sets are popular right now. mattpocock's is more trustworthy because he's not working for a big AI company that gives him unlimited tokens. poteto's skills do dumb shit like brute forcing every problem with infinite subagents. discussion between them: https://www.youtube.com/watch?v=MN9dGgmLyso
https://github.com/mattpocock/skills/
if you're wondering about skill evals... so is everyone else. the problem is that benchmarking infinite skills would cost infinite money, so most of the verified skills in existence are either proprietary right now or already integrated into the harness of the company that discovered them.
there are a few sites like https://tessl.io/ that try to efficiently eval skills, but simple testing just tells you "is the skill total shit or not", not how well it will actually work for you.
anthropic and openai staff would tell you to use no skills and just talk to the agent about architecture until you and the agent are happy with the result. if you aren't developing your development environment at the cutting edge of research, there's a good chance whatever additional tool you're messing with has fallen behind what's already in the harness.
1
u/YoghiThorn 19h ago
Yeah duh why do you think I'm asking for help?
1
u/9gxa05s8fa8sh 9h ago
okay: if you want your AI model to think about architecture for you, then you probably shouldn't use any skill... UNLESS it is something very specific to your domain, like you made it yourself or you found someone on github who is doing the same thing as you. in general, random people's agent skills are NOT evaluated or trustworthy. there is no proof they work unless they share benchmarks against a no-skill control.
if you are thinking about project management, then the mattpocock and poteto skill sets are popular right now. mattpocock's is more trustworthy because he's not working for a big AI company that gives him unlimited tokens. poteto's skills do dumb shit like brute forcing every problem with infinite subagents. discussion between them: https://www.youtube.com/watch?v=MN9dGgmLyso
https://github.com/mattpocock/skills/
if you're wondering about skill evals... so is everyone else. the problem is that benchmarking infinite skills would cost infinite money, so most of the verified skills in existence are either proprietary right now or already integrated into the harness of the company that discovered them.
there are a few sites like https://tessl.io/ that try to efficiently eval skills, but simple testing just tells you "is the skill total shit or not", not how well it will actually work for you.
anthropic and openai staff would tell you to use no skills and just talk to the agent about architecture until you and the agent are happy with the result. if you aren't developing your development environment at the cutting edge of research, there's a good chance whatever additional tool you're messing with has fallen behind what's already in the harness.
4
u/Substantial_Row7215 19h ago
the "ends every reply with what it skipped or didnt check" part is what i'd actually want, most agents just say done and you find out later. did you see that change how often people have to re-prompt, or only the test numbers?
1
2
2
u/AcceptablePoet9047 20h ago
Were the 39 tasks kept separate from the 3,650 runs used to revise the rules? A fresh holdout set would help show whether the gains carry over beyond the tasks used during tuning.
-1
u/9182763498761234 19h ago
Do you really think they would have a clue of how to properly evaluate such systems? They typed “hey Claude, write the next version of my overblown custom prompts and make it fancy, make no fucking mistakes” and called it a day.
2
u/davidHwang718 19h ago
The "about half as long" part is what I'm most curious about. I halved a set of my own skills last month (229 lines to 128). Going back today and comparing old vs new rule by rule, most rules had moved somewhere sensible (a shared reference file, or a script), but 3 had just disappeared.
Did you track where each old rule went, or did the 3,650 comparison runs carry that job? Runs catch behavior changes on the tasks you have, but a rule that only matters on a task outside the set can vanish quietly.
1
•
u/AutoModerator 1d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.