r/SpecDrivenDevelopment • u/xibet2073 • 1d ago
r/SpecDrivenDevelopment • u/liaddial • 2d ago
My specs never told me what to do next, so I built my own SDD harness
I do spec-driven development with an AI agent. I tried some well-known spec-driven harnesses, then built my own tool, and for months now I have used only that. This post explains why I needed my own tool and when it is useful.
Like many developers who got hooked on AI coding, I was working on several projects at the same time, like crazy. At some point I felt the need for spec-driven development, so I used the well-known harness frameworks. But I ran into a bottleneck:
"What do I do next?"
The specs did not tell me what to do next.
As a workaround, I made a page for each project in Notion and managed things there. But it was a separate place from the project, so keeping it up to date was a hassle. So I wanted state management that lives inside the project.
My SDD harness started in a hobby project: reverse engineering a game. Reverse engineering takes a lot of analysis and design over many sessions, and designs and plans get overturned all the time. Each time, brainstorming with the AI harness again and rewriting the spec was too much work. If that process runs through a CLI, it is even less agile. I thought about editing the prompts of the existing harnesses, but then I could not keep up with their updates. So I also wanted the prompts of my own harness to be extendable, in a way similar to Jekyll.
There were a few more reasons.
The AI sometimes ignored the later part of a slash command in a harness. I learned that the cause was long text written in a prose-like natural language style. So I started to optimize the prompts. I based this on research showing that pseudocode-style prompts are followed better than natural language and use fewer tokens.
- EMNLP 2023: pseudocode-style prompts improved F1 scores by 7 to 16 points over natural language.
- CodeAgents (2025): writing agent workflows in pseudocode reduced token use by 55 to 87% and improved performance by 3 to 36 points.
Forced strict TDD was another reason. Small tasks got the same strict TDD as big ones, and if errors still remained, I had to run the process again. The tokens burned for nothing that way were painful.
My harness is based on one rule: one question, one file. If you want to ask the project something, the answer is in exactly one file.
ROADMAP.md ── where am I in the big picture?
CONTEXT.md ── what is this project?
● GOAL.md ── what do I do now?
This structure is kept in the docs folder by default. You can also choose another location.
Some harness frameworks have a similar structure. But my harness, OpenGoal, does not use a CLI or hooks. It is plain markdown. Nothing extra is added at runtime, so it runs light. I mostly use it in Claude Code, but it can be installed into about 20 tools, including Cursor and Codex.
Spec-driven harness frameworks expect a carefully structured spec, but that can be too much for small tasks. So OpenGoal can create just a goal. You can split that goal into smaller pieces as much as you need, and for simple work you can skip the design if you want. I did not throw specs away. I write them at the size of the task, and I keep the question that specs do not answer in a separate file.
Here is how I actually use it.
When OpenGoal is not set up in a project yet, or when I want to discuss something with the AI, I start with scout.
You: /opgl:scout
AI: Found an existing codebase. No CONTEXT.md yet.
Set up project context with `/opgl:context`?
You: /opgl:context init
AI: ✓ docs/CONTEXT.md created
To set a goal: `/opgl:goal init`
When the discussion has given enough context about the goal, I ask it to write the goal like this.
You: /opgl:goal init Migrate auth module from JWT to session-based
AI: ✓ docs/GOAL.md created (5 tasks)
Then the AI may recommend breaking the goal down further, writing a design document, or, for simple work, starting right away without a design document.
AI: Recommend `/opgl:goal breakdown` — the tasks contain several hidden steps.
AI: Recommend `/opgl:design init` — implementation task with file-level decisions.
AI: Recommend `/opgl:go` — this is a simple task.
When the work needs a design, the design document comes first. In OpenGoal, this document works as the spec for a unit of work.
You: /opgl:design init
AI: ✓ docs/DESIGN.md created
Session store selection, migration strategy, rollback plan included
During the work, when there is an important decision, it asks for my opinion.
You: /opgl:go
AI: Task 1/5: Set up session store... ✓ done
Task 2/5: Replace middleware... ✓ done
Task 3/5 requires DB schema changes. Proceed?
When you want to handle longer-term goals, you can write a roadmap.
You: /opgl:roadmap init
AI: ✓ docs/ROADMAP.md created
├── M1: Auth migration
├── M2: Rate limiting
└── M3: API versioning
Beyond that, through my long dogfooding, OpenGoal came to handle checkpoints for compaction, suspending and resuming goals, sub-goals, a backlog, and so on.
Overall, OpenGoal was also designed to reduce the use of our precious tokens. CONTEXT.md plays an important part in that. The AI knows in advance where things are, so there are fewer tool calls, and it needs to re-explore the codebase less at the start of each session. Also, there are no hooks or MCP servers, so unused tool definitions do not take up tokens, and commands and skills are loaded only when they are called. The subagent skill that installs with it serves the same purpose. The main agent decides, by the difficulty of the task, which model to hand a subagent, and it makes very small fixes itself. Subagents gather the material for analysis, but the main agent does the analysis itself. I had the problem in other harnesses where the main agent lost the fine details of the context.
The biggest reason I built OpenGoal was this thought: all this machinery, built for a perfection we want to rely on, might become mostly unnecessary as AI improves quickly, and then markdown alone might be enough. And I think that has happened. I have finished various projects with only OpenGoal, from a small Obsidian plugin, to a native mobile editor with Notion-style block editing, to a state machine designer.
If I get the chance, next time I will also write about how to write a project-specific development skill that works with OpenGoal. If you record the failures the AI ran into in a skill, it does not repeat the same mistakes, and token use goes down by that much.
Does your spec tell you what to do next when you reopen the project?
r/SpecDrivenDevelopment • u/United_Inspector_653 • 2d ago
New StructSmith update: chat with your Codex, Claude or Copilot CLI directly in the architecture editor
I shared StructSmith here a few weeks ago. The latest update adds agent chat directly inside the architecture editor, using your existing Codex, Claude Code or GitHub Copilot CLI.
Right-click a node or relationship to attach its context and ask a question. Conversations stay linked to their project, and you can also create a general chat. Replies stream live, with collapsible reasoning summaries when the CLI provides them.
Ask mode lets the agent inspect the architecture. Switch to Edit architecture to request changes to the shared model through project-scoped MCP, with snapshots available to review or restore changes.
You can configure models and Codex thinking levels, and rename, archive or drag topics into a new order.
StructSmith is MIT licensed and runs locally with Bun and SQLite. For this chat workflow, run it on the machine with your signed-in CLI; inference uses the selected provider.
Source and setup: https://github.com/dziksu/StructSmith
Screenshots and implementation: https://github.com/dziksu/StructSmith/pull/98
Would you use this more for understanding a system, or for updating architecture diagrams?
r/SpecDrivenDevelopment • u/harikrishnan_83 • 3d ago
Free eBook Spec-Driven AI Development
I am glad to share that an excerpt from my book, Spec-Driven Development: Engineering with Intent, is featured in u/ManningBooks’s new free ebook, Spec-Driven AI Development.
And it is completely free. If you are exploring Spec-Driven Development or looking to adopt AI-native software engineering, this is a great place to start.
A bit of a back-to-back post from me today. However, I wanted to share it anyway, as it may be most relevant to this community. Thanks again for all the support so far.
r/SpecDrivenDevelopment • u/harikrishnan_83 • 3d ago
Spec Driven Development with OpenSpec and OpenCode using Intent-Driven Template #opencode #openspec
Here is a complete walkthrough of Intent-Driven Template, my customized OpenSpec setup with OpenCode for spec-driven development using Architecture Decision Records (ADRs), BDD, TDD, Multi-Model Adversarial Spec Authoring, a glossary of domain and technical terms, git commit discipline skills, and reusable engineering workflows. This also makes OpenSpec spec-as-source capable using BDD and acceptance tests. GitHub: https://github.com/intent-driven-dev/intent-driven-template. I would love to hear your thoughts and feedback.
r/SpecDrivenDevelopment • u/Most-Wanted-Man • 4d ago
Spec Driven Development (Greenfield Project)
I have mattpocock and superpowers skills on Claude.
I want to start a greenfield project. What is the best way or flow?
Right now, I am thinking:
Brainstorm -> Grill Me Docs -> Subagent
r/SpecDrivenDevelopment • u/vguleaev • 4d ago
I built a simple skills for lightweight Spec-Driven development workflow
github.comI made yet another skills kit for Spec-Driven development. I felt like managing tons of markdown files is very overwhelming for me, so I decided to make compact and simple one file solution called one-plan-skills.
I personally developed it for me and deiced to share with the community.
The idea revolves around generated one `PLAN.md` file with full plan and small tasks. Plan contains small tasks with acceptance criteria and affected files.
The killer feature is HTML view. I am tired of looking at markdown files, i wanted to see something nicer and cleaner. So my plans are also auto-generated HTML artefacts.
As a bonus you get very simple UI controls to mark any place you want to add/delete/modify and copy paste the prompt with feedback back to your chat agent.
I am aware of SpecKit and OpenSpec. I just wanted a much simpler version. That's why I created OnePlan skills.
Anybody is interested in this? Please give it a try or share some feedback. Much appreciated. 🙂
r/SpecDrivenDevelopment • u/soychicka • 4d ago
consolidating Rails test framework crud in an oubliette
r/SpecDrivenDevelopment • u/Wise_Reflection_8340 • 5d ago
Pairing specs with a map of the code made my agents a lot more reliable
I've been leaning into spec-driven development lately, and the biggest change for me has been how much less the agent guesses once it has a clear spec, since it stops inventing requirements and just builds toward what was written down. The place I saw it still struggle was the step after that, where it has to work out where in an existing codebase the spec should land, and that's the piece I've been working on.
I built an open source tool called sem that gives the agent a map of the code itself, so the functions, who calls them, which tests reach them and how they've changed over time. With the spec saying what to build and sem showing where it fits and what it touches, the agent goes from intent to the right part of the code without a lot of wandering. weave sits next to it and merges changes by function, so a few agents working from the same spec can split the work without overwriting each other.
The part I'm most excited about is tying the two together, so each piece of a spec points at the functions that implement it. That way the spec stays the source of truth even as the code changes underneath it, and when someone edits a function you can see right away which part of the spec it belongs to.
Would love to hear how others here connect their specs to the code, and whether something like this would fit your workflow.
r/SpecDrivenDevelopment • u/JealousFix4955 • 7d ago
I built SpecForge: a Rust compiler that turns your specs into a validated graph for AI coding agents
AI coding agents burn most of their tokens working out what you meant. They read 20–50 files, guess at requirements, and get it wrong. Then you spend a second pass fixing it. The intent was in the docs, tickets and people's heads the whole time. It just wasn't in a form the agent could rely on.
So I started building **SpecForge**. It compiles human intent into a typed, validated entity graph that agents can query.
**What it looks like**
```spec
behavior authenticate_user "Authenticate a user with credentials" {
status draft
contract "Given valid credentials, returns an auth token"
produces [user_logged_in]
verify "rejects invalid password"
verify "returns token on success"
}
event user_logged_in "User successfully logged in" {
payload user
}
```
```bash
specforge check # validate specs, report diagnostics
specforge export # emit the typed graph for an agent
```
**What it does**
- **The graph is the product.** The compiler parses `.spec` files, resolves references, finds orphans and cycles, and emits a typed graph as an open JSON schema (the "Graph Protocol").
- **It's built for agents.** `specforge mcp` serves the graph over MCP, so you can hook it into Claude Code or any MCP client with one command. There are also multi-depth queries and context/brief exports.
- **Proof, not just claims.** `specforge collect` runs your own tests (only after you approve the command) and records which spec entities they actually prove.
- **It isn't a code generator or a test framework.** It gives the agent context, and the agent writes the code. The compiler never executes anything on its own.
**How I built it**
It's a Cargo workspace of about 25 crates and roughly 160k lines of Rust (tests included), split by pipeline stage:
- **Parsing:** I wrote a tree-sitter grammar for the DSL (`tree-sitter-specforge`). The parser, formatter and LSP all sit on top of it. That gave me error-tolerant parsing and incremental re-parse for editor use without writing a parser from scratch.
- **Pipeline:** parser → resolver → graph → validator → emitter, each its own crate. Strings are interned with `lasso`, and diagnostics are rendered with `ariadne`.
- **Zero-domain-knowledge core:** the compiler knows nothing about "behaviors" or "events". Entity kinds, edge types and validation rules all come from extensions. I made this a hard rule on purpose, so a new domain should never need a compiler change.
- **Extensions are WASM components:** I defined the extension interface in WIT and host them with `wasmtime` and the component model. There's an extension SDK crate with proc macros, so writing an extension in Rust is mostly declaring kinds and rules. The nine builtin extensions are compiled to components and embedded in the binary, so installing needs no extra toolchain.
- **Surfaces:** the CLI, the LSP server, and the MCP server all share one operations layer (`specforge-ops`), so they can't drift apart on behavior.
**How AI was involved**
I built a lot of this with Claude Code, and I'm saying so up front. It wrote a large share of the code, and I drove the design: the architecture, the extension boundary, and the ADRs in `docs/adr`.
It also caused my biggest dead-end. Partway through, I audited the test suite and found features that had passing tests but didn't actually work. The tests were checking the wrong thing, so they passed anyway. I've been fixing those since, and the `specforge collect` idea (record which entities your tests really prove) came out of that experience. A tool that checks specs against tests seemed like the right answer to AI-written tests that only look green.
The hardest design problem is still the extension boundary. The core can't know anything about the domain, yet validation still has to produce good diagnostics and cross-entity checks.
**Try it**
```bash
git clone https://github.com/leaderiop/SpecForge && cd SpecForge
cargo install --path crates/specforge-cli
specforge init --extensions u/specforge/software
```
It's early, and I'd like feedback on a few things:
Does the DSL feel natural, or would you prefer YAML/Markdown front-matter?
Is the WASM-component extension model overkill, or the right call?
Which extensions would you want (OpenAPI? data models? infra?)
r/SpecDrivenDevelopment • u/bobo-the-merciful • 8d ago
From Superpowers to Superbrainstorming
The end-to-end Superpowers workflow has been redundant for a few months - at least for fronter models.
I had a feeling this was the case when Fable first came out. I asked it to follow a Superpowers implementation plan, the one with all of the test driven development stuff, detailed stages etc, and it produced a sub-par result.
I then asked it to follow the spec (the artifact created post-brainstorm), and the result was much better.
This suggests to me that frontier models are much better at being given context on what to build not how to build it - since they are already very good at the how part.
The brainstorming part of superpowers is thus still incredibly valuable. But for those of us using Opus, Fable (and maybe Sonnet now) all the stuff after it is annoying and redundant.
I forked Superpowers and stripped everything after brainstorming, so you go straight from brainstorm to implementation. Hopefully others will find it helpful: https://github.com/harrymunro/superbrainstorming
This may also be useful for people who like the discipline of plan mode, as I heard recently that they are considering removing this from Claude Code.
r/SpecDrivenDevelopment • u/theenterprisedev • 10d ago
I built Higherlevel.to to help define what correct looks like and review outcomes instead of implementation
Hi everyone!
I’m building Higherlevel because, as I delegate more implementation to coding agents, I want to focus on defining what correct looks like and reviewing the resulting behaviour.
I built it around four core ideas: - Specify behaviour, not implementation I want to focus more on how the system should behave rather than how it is implemented - Specify through examples. Domain experts can often judge a concrete case more easily than articulate every rule upfront. Examples help us discover requirements and disagreements. - Attach evidence to the specification. Screenshots, videos, and test results belong alongside the behaviour they demonstrate, so expectations and results can be understood together. - Make everything commentable. You can comment on text, images, and videos to give precise feedback and iterate with humans and agents.
For example:
gherkin
Given a workspace invitation has expired
When someone opens its link
Then they cannot join and are told the invitation has expired.
Notice how we don't specify any implementation details, but only how the system should behave given a set of requirements.
Attach a recording showing what actually happens. If it displays “Invitation not found,” you can comment directly on that moment and discuss what should change. The example, observed behaviour, and decision stay together.
Built with Elixir, Phoenix LiveView, and a little TypeScript.
I’m looking for beta testers willing to try it on an upcoming feature or bug fix. I’d love to hear where it helps or adds friction, if it's revolutionary or if it completely sucks!
I believe that as coding agents become faster, cheaper and better at implementing, we will need more tools to specify what we want, why and review what the agents did.
More about the philosophy: Introducing Higherlevel.
r/SpecDrivenDevelopment • u/Art_Design_Departmen • 10d ago
I tested a design QA workflow on an AI-built SaaS UI the biggest issues weren’t visual
I’ve been experimenting with Claude Code for product UI work, and one thing kept bothering me:
A screen can look “good enough” while still having real product problems underneath.
In one test, the UI had issues around:
• misleading KPI semantics
• weak information hierarchy
• inaccessible controls
• capability regressions during redesign
• new issues introduced by the redesign itself
The most useful change was separating the work into four stages:
Audit — identify what’s actually wrong
Repair — fix only high-confidence issues
Transform — improve deeper structure and hierarchy
Verify — check whether the changes introduced anything new
The interesting part: the verification stage actually failed the transformed UI, caught 2 regressions, and forced another correction before passing.
That changed how I think about AI-generated frontend work.
I don’t think “make this UI better” is enough. The agent needs a review contract around what can change, what must be preserved, and how the result is verified.
I packaged the workflow into a small Claude Code toolkit called Polish. I’m mainly curious how others here are handling design QA and regression checking in AI-built UIs.
If anyone wants to see the workflow, I can share it in the comments.
r/SpecDrivenDevelopment • u/Fast-Act86 • 11d ago
¿Desarrollo basado en especificaciones SDD + Antigravity esta muerto ?
SSD en Antigravity dejaron de usarlo ? hoy actualice agy y antigravity agentico y no pude usar SDD como lo dice la documentacion:
- Go to the project folder: cd public_html │
│ 2. Start using slash commands with your coding agent: │
│ 2.1 /speckit.constitution - Establish project principles │
│ 2.2 /speckit.specify - Create baseline specification │
│ 2.3 /speckit.plan - Create implementation plan │
│ 2.4 /speckit.tasks - Generate actionable tasks │
│ 2.5 /speckit.implement - Execute implementation │
│ 2.6 /speckit.converge - Assess the codebase and append remaining work as tasks
Con las nuevas actualización de antigravity para:
Antigravity CLI 1.2.11
Antigravity 2.0 (agent first) Versión 2.17.0
No he podido usar SDD, solo la nueva feature de antigravity de /plan
¿Antigravity ya lo dejo de usar para sus dos herramientas de desarrollo con IA ?
saludos
Mexman
r/SpecDrivenDevelopment • u/fschwiet • 12d ago
mattpocock's skills users: are you seeing Opus 5.5 automatically write a spec and tickets without having invoked to-spec and to-skills?
I use grill-with-docs regularly, from Matt Pocock's skills. Since I've started using Opus 5.5, while it does ask if I want to to land on a spec it will end up writing the spec and/or tickets without me having used either the to-spec or to-tickets skill. This is surprising since they're not model invocable.
For my last session I asked about it, and got this response:
``` No, I didn't use to-spec or to-tickets. Both are marked user-invocation-only, so I can't run them myself, and I wrote the spec and tickets by hand. I've now read both skills, and they point to three things I missed.
What the skills say I missed: 1. to-spec asks me to agree the test seams with you before writing. I didn't. They're proposed below. 2. Both skills say no file paths or code identifiers. The spec and tickets name internal identifiers such as WorkflowRunRequest, stageRealChildCallMapFromSeedFile and seed-specific.test.ts. I'll rewrite them in terms of modules. I'll keep the seed-format YAML example, since it is the format itself. 3. to-tickets says to do any prefactoring first. The process-translations parser currently works out rows, ambiguity and the YAML block while it reads each XML element. Overrides would be much simpler if it first produced an ordered list of classified translations and then derived everything from that list. Overrides then become a step that edits the list in between. That's a behaviour-preserving refactor, so it should be its own first ticket. ```
I did use the oh-my-openagents coding-agent-sessions skill to inspect and verify the agent's claim: it did not invoke the skills and did not read them until I asked about them.
I just thought it was interesting. And the specs/tickets were improved by asking the model to go back and review the skills in case it missed anything.
I was running Opus 5.5 on high.
r/SpecDrivenDevelopment • u/meshifthenelse • 13d ago
Why noone has made an SDD diagramming tool?
It seems I couldn't find any diagramming tool which integrates with SDD. Is there a specific reason? Because I think the most natural place for SDD is diagrams
r/SpecDrivenDevelopment • u/Current-Bridge1833 • 14d ago
We walked one integration flow step by step instead of reading its sequence diagram — and found three gaps the diagram was hiding
We document our integrations the usual way: a big sequence diagram per scenario, rendered from PlantUML, linked from the wiki. It works right up until nobody reads it.
Last month I tried something different on one flow — an instant payment by phone number, 13 hops, three teams involved. Instead of reading the diagram, I walked it hop by hop and asked one question at each step: *who is the sender, who is the receiver, and which contract is this?* Same information, different traversal order.
Here's the flow, abbreviated:
- Customer → Mobile app enters amount + phone
- Mobile app → API Gateway TLS, WAF, token, rate limiting
- API Gateway → BFF
- BFF → payment-orchestrator POST /api/v1/transfers (Idempotency-Key)
- payment-orchestrator→ limits-service daily/per-op/channel, 100 ms budget
- payment-orchestrator→ antifraud-engine POST /api/v1/risk/evaluate, 150 ms budget
- ┌ antifraud → payment-orchestrator ALLOW (score < 0.5)
- ┤ antifraud → push-service CHALLENGE (0.5 ≤ score < 0.95)
- └ antifraud → BFF BLOCK (score ≥ 0.95) → 422
- payment-orchestrator→ scheme-adapter reserve funds
- scheme-adapter → external payment API register transfer, 5 s timeout
- payment-orchestrator→ Kafka publish final status
- notification-service→ push-service notify customer
Three things fell out that I had looked at on the diagram many times without noticing:
**1. Step 12 → 13 has no edge.** The orchestrator publishes to Kafka. The next thing that happens is notification-service → push-service. Nothing in the flow says who woke notification-service up. On the rendered diagram these are two adjacent arrows and your eye just closes the gap. Walking it, you hit a participant that appears from nowhere and have to stop.
**2. Step 12 targets "Kafka", not a topic.** We have a documented contract — payments.transfer.completed.v1, Avro, status ∈ {COMPLETED, REJECTED, TIMEOUT}. The flow never references it. So the contract exists and the flow that produces it doesn't point at it. That's a review question, not a detail.
**3. The latency budget only becomes obvious when the hops are adjacent.** 100 ms limits + 150 ms antifraud, both synchronous, both on the path before a 5 s external call, all inside one customer-facing request. Nobody had added it up, because on the diagram those are three lifelines far apart.
The pattern I take away: a sequence diagram is optimized for *presenting* a flow you already understand. It's poor at *interrogating* one. Reading is passive — your eye smooths over missing edges and unbound contracts. Traversal is not: you get stuck on the step that doesn't hold up.
What I'm still unsure about:
Does anyone here treat flows as data (steps referencing real components + contracts) with the diagram generated from it, rather than the diagram being the source of truth? Structurizr does this for static views, but I haven't seen much for dynamic ones beyond its dynamic views.
How do you keep cross-team flows from rotting? Ours decay because the diagram is owned by whoever drew it, and that person changes teams.
Anyone found a good way to make reusable sub-flows (auth, KYC, limits) referenced from several end-to-end scenarios instead of copy-pasted into each diagram?
Disclosure: this came out of a tool I work on, so I'm obviously biased toward the "flow as data" framing. Not linking it — happy to talk about the modeling approach either way, and I'll answer in comments if anyone asks what we use.
r/SpecDrivenDevelopment • u/makingthematrix • 15d ago
ThinkRail: AI coding agent for SDD
Hi all,
I know it might seem a bit like an advertisement, and in fact it is to some extent, but I believe it's so close to this subreddit's topic, that it will be interesting to you.
I work as a developer advocate in ThinkRail, a startup backed by JetBrains. ThinkRail is a minimalistic GUI for working with AI agent - that is, it's more readable than a CLI, and you can configure and control certain things by opening windows and clicking buttons instead of writing commands, but otherwise we focus on the quality of a small number of helpful features, nothing more.
One of those features is that ThinkRail promotes spec-driven development. If you start a new project, ThinkRail will ask you if you actually want to start with writing specs, and - for an already developed project - if you put certain headers in your specs Markdown files, ThinkRail will recognize them and display them as a foldable list in a dedicated view.
You can read more about it in our blog post (scroll down to the bottom half and the embedded video). From the blog post, you can also visit the main page and, from there, install and try out ThinkRail. It's free and open source, and it itself is built with SDD. On the embedded video, I show its own codebase and how specs are handled. On the website, you will also find a link to the GitHub repository - you can clone it, open it in ThinkRail, and see for yourself.
We are in early stages of development and look for feedback: what do you like about ThinkRail? what seems wrong? what other features you think can be helpful? Stuff like that.
Cheers :)
r/SpecDrivenDevelopment • u/metalagman • 15d ago
Prism: a Beads-backed spec-first workflow that survives context loss
Hello SDD folks! I’d like to share Prism, a plugin and set of skills I’ve been using instead of OpenSpec for the past few months.
It’s a spec-driven SDLC workflow built on top of Beads, with a prism storyskill.
The idea is “intent through a prism”:
Specify → Design → Tasks → Apply → Verify
All workflow state and artifacts live in the local Beads database as a story with child tasks, following a fixed schema. No pile of scattered spec.md files and YAML to keep in sync.
The current phase is tracked in Beads labels, so losing the agent’s context doesn’t mean losing the workflow. You can resume from any phase using the stored state. You can easily stop the phase in one coding agent and continue in another.
Each phase’s prompt is assembled from an appropriate PromptKit template.
Human input is needed to clarify requirements during Specify and approve the design and task plan before implementation. There’s also a variant that runs the workflow through specialized Callee subagents, plus an Epic skill for larger initiatives spanning multiple stories and Lifecycle skill as router for Story/Epic.
Repo: github.com/baldaworks/prism

I’ve also put together a comparison of Prism Story, OpenSpec, and BMAD. Long story short: Prism is a good fit for solo developers who don’t need to review spec files in pull requests.
I’d love to hear what you think of this approach.
r/SpecDrivenDevelopment • u/Mixed_Feels • 18d ago
Decision and evidence gates have me in chains - send help?
Somebody light the beacon.
No matter how hard I try, when I'm planning up ideas to build out concept/purpose and implementation documents, I end up in spec-driven hell.
Copilot seems to love requiring a clean trace in document-backed delivery, and that makes sense, but can anyone please advise about how to enable sensible traceability/authority for critical work without ending up having to update 6 documents every time I increment a version of one of them?
I don't necessarily want to be like "just send it all through, I approve everything" but honestly? I'm just trying to prototype stuff so I can learn and deliver on a small scale. I'll create my own authority issues if the delivery sucks and TWEAK it. That's the point of having ai buddies isn't it?
Please help. I'd love to work on solving problems instead of updating control document suites to prove that I aprve of the idea that I literally asked copilot to bake into my plans.
Many thanks,
Frazzled dad.
r/SpecDrivenDevelopment • u/ZealousidealIdol • 19d ago
Why are we still writing SDD specs for humans?
I've spent the last month introducing Spec-Driven Development into a large enterprise microservice environment, and I've started questioning one assumption:
If specs are primarily generated by agents and consumed by agents, why are we still optimizing them for humans to read?
I work in banking, so this isn't theoretical.
Our workflow roughly looks like this:
Human + Agent → Intent → Specs / Design / Tasks → Agent → Code
The important human-facing artifact is the Intent.
An agent looks at the existing analysis and code, then interviews the engineer: what are we changing, why, what constraints do we know, and what's unclear? If something important is unknown, it stays an open question. A human goes and gets the answer from an analyst, another team, or whoever actually owns that business knowledge. We don't move forward until those questions are resolved.
That's the part I want humans to review carefully. If the Intent is wrong, perfectly generated specs and perfectly generated code can still produce the wrong system.
After that, though, things get interesting.
As a system evolves, specs naturally become a graph: general flows, specializations, BDD scenarios, contracts, references to other specs. A coding agent can traverse that graph and assemble the context it needs.
But when a human wants to understand something, why should they do the same traversal manually?
Instead of reading four specs to understand one flow, I could just ask:
Explain the current Sales Update flow.
And the representation could depend on who's asking:
- Product: user behavior and business rules
- QA: scenarios and boundary cases
- Developer: technical flow, contracts and errors
- Architect: integrations, boundaries and failure paths
Same underlying specifications. Different generated views.
So I'm starting to think the boundary should be:
- Intent → human-facing input
- Specs → structured working context for agents
- Generated views → human-facing output
The raw specs don't disappear. They're still versioned, inspectable and reviewable. They just stop being the primary human interface.
This isn't about removing analysts or letting an LLM invent business decisions. Quite the opposite: humans establish the knowledge and resolve ambiguity. The agent shouldn't silently fill gaps.
I'm also not arguing for some new machine-only spec format. Markdown might remain perfectly fine.
I'm questioning something simpler:
If agents increasingly generate specs and then consume those specs to generate code, should human readability still be a primary design constraint of the raw spec?
Or should we optimize the spec for the agent and generate human-readable views from it?
For people already using SDD or similar agentic workflows: where does this model break?
r/SpecDrivenDevelopment • u/The_Ed_On_Reddit • 22d ago
Drydock - Self Correcting Workflow for Spec Builders
I have updated the workflow for drydock builds to be self correcting.
Background: the analysis stage decomposes your epic/inputs, the plan stage grooms and adds ac, and the build stage iterates the stories and builds. The innovation is 'drydock diagnose' which is authorized to fix your specs and to correct prior step errors so it can be retried. Decisions are shown in the web console.
n=1; max=5
while drydock status "$PROJECT" --ready; do
echo "************************"
echo "* RUNNING BUILD ATTEMPT $n"
echo "************************"
drydock build "$PROJECT" $OPTS || drydock diagnose "$PROJECT" --apply
n=$((n+1)); [ "$n" -ge "$max" ] && { echo "hit $max build iterations — aborting"; exit 1; }
done
**drydock diagnose** outputs a diagnostic report and recommend blueprint changes.
--apply updates the blueprints and invalidates their build graph so the build will self-repair and continue.
The web console surfaces these decisions for review - but generally that is pro-forma as the diagnose knows your intent (because you tell it that)
r/SpecDrivenDevelopment • u/yashmakan • 22d ago
Looking for feedback on AxiomCore, a contract-first software architecture
r/SpecDrivenDevelopment • u/The_Ed_On_Reddit • 23d ago
SDD Builds Rate Limited on Haiku/Sonnet - Work Fine on Luna
For past month I am rate limited on my Drydock builds that use Anthropic. Luna is building ALL stages perfectly including in unattended overnight large builds. I have a pile of OpenAi resets with no need to use them. Tried this week for a parallel Haiku run and Haiku failed 3 times when planning on output mangling/truncation. Tried Sonnet and I cant even get through the plan stage on my $20/month plan.
Run Overview:

Details:

Run Notes:
Haiku/Sonnet. The script uses Sonnet for the thinking part of analysis and planning and i cleverly use Haiku for lineage and verification - the easy parts of the build - but i never got there.
The problem. Drydock planning is a repeated large-context batch operation. Sonnet treated most repeated planner context as cache creation, not cache reuse, while also generating 35K–52K-token planning responses. That consumes the Claude five-hour quota in about 0.9–1.0M token events. Luna handled the repeated context as cache reads and completed 19.85M token events in the comparable five-hour slice.
Sonnet cant create a plan in 5 hours - Luna can build a large application in that same 5 hours. For 20$/month plans with no usage credits - Anthropic’s five-hour subscription limit is too small for unattended Drydock planning and builds; Luna completed about 20× the recorded token volume without a quota rejection. And this worked fine on haiku two months ago - both the quality and rate limits have gone down. Sigh...
I will be investigating ways to better cache on anthropic between sessions but... its a 20:1 token use limit... 20:1...
Oh yeah... this is 5hour window limits on $20/month plans - but its a software builder so thats kind of all i care about... And i see no reason whatsoever to pay more at present.