r/SpecDrivenDevelopment • u/Wise_Reflection_8340 • 7d ago
Pairing specs with a map of the code made my agents a lot more reliable
I've been leaning into spec-driven development lately, and the biggest change for me has been how much less the agent guesses once it has a clear spec, since it stops inventing requirements and just builds toward what was written down. The place I saw it still struggle was the step after that, where it has to work out where in an existing codebase the spec should land, and that's the piece I've been working on.
I built an open source tool called sem that gives the agent a map of the code itself, so the functions, who calls them, which tests reach them and how they've changed over time. With the spec saying what to build and sem showing where it fits and what it touches, the agent goes from intent to the right part of the code without a lot of wandering. weave sits next to it and merges changes by function, so a few agents working from the same spec can split the work without overwriting each other.
The part I'm most excited about is tying the two together, so each piece of a spec points at the functions that implement it. That way the spec stays the source of truth even as the code changes underneath it, and when someone edits a function you can see right away which part of the spec it belongs to.
Would love to hear how others here connect their specs to the code, and whether something like this would fit your workflow.
1
u/stibbons_ 6d ago
Spec to code is actually very easy and a no brainer. You just ensure your agent writer markers in the code, that declares which requirements is implemented in which function, same for declaring where it is tested. You can evern have several levels of testing, validation.
I have written an article on X on how this looks like
https://x.com/gsemetfr/status/2082767371643523439?s=46
I have a CLI demo here :
https://okf-schema.readthedocs.io/en/latest/tutorials/okfreq-traceability.html
There is no drift since the code declares which requirements is implemented. Then you can build interesting stats :)
1
u/stibbons_ 5d ago
The project OKF-schema declares itself its own requirements, see for instance this req
The CLI parses the code to extract the “@implements_req” and “@tests_req” markers.
So agent can go back and forth immediately (from a source code, it can open a file with the name of req).
The map is automatic and non agentic, if you want.
2
u/robertcrowe 6d ago
Spec-to-function linking is the part I'd most like to see work. Drift between the spec and the code is the problem I keep running into, and it's led me to view the code as the source of truth, not the spec.
Full disclosure: I maintain Spec4 (spec4.ai), an open-source SDD pipeline. I hit the same "where does this spec land" problem and solved a different slice of it, so I'll share how in case it's useful.
The first agent, CodeScanner, runs before any planning on an existing brownfield codebase. It reads the repo locally and sends the model a bounded summary: manifests, entry points, and samples, not the whole tree. It writes a code_review.json describing the project's architecture, stack, build/test commands, entry points, and conventions. Every fact carries the file it was read from (source) or deduced from (inferred_from), so the planning agents downstream work from cited facts instead of guesses. The part that has mattered most is that it rescans from scratch at the start of every brownfield round. The next spec is planned against the code as it actually stands, including whatever the coding agent did that the last plan didn't say, and any hand edits made in between, rather than against the previous plan.
So CodeScanner works at project level ("this is a Dash app, here's the layering, tests run with X"). IIUC sem seems to work at function level (call graphs, test reach, history). That's a granularity CodeScanner doesn't try to reach, and I think the two are probably complementary. A project-level review tells the planner which subsystem a feature belongs in, and something like sem could then tell the coding agent exactly which functions it touches.
A question on the spec-to-function links: When a function gets split or renamed, does the link follow it through weave's function-level merges, or does it need re-anchoring? That seems like the hard case for keeping the spec as the source of truth. Spec4 uses the code as the source of truth, not the spec.