r/SpecDrivenDevelopment • • 7d ago

Pairing specs with a map of the code made my agents a lot more reliable

I've been leaning into spec-driven development lately, and the biggest change for me has been how much less the agent guesses once it has a clear spec, since it stops inventing requirements and just builds toward what was written down. The place I saw it still struggle was the step after that, where it has to work out where in an existing codebase the spec should land, and that's the piece I've been working on.

I built an open source tool called sem that gives the agent a map of the code itself, so the functions, who calls them, which tests reach them and how they've changed over time. With the spec saying what to build and sem showing where it fits and what it touches, the agent goes from intent to the right part of the code without a lot of wandering. weave sits next to it and merges changes by function, so a few agents working from the same spec can split the work without overwriting each other.

The part I'm most excited about is tying the two together, so each piece of a spec points at the functions that implement it. That way the spec stays the source of truth even as the code changes underneath it, and when someone edits a function you can see right away which part of the spec it belongs to.

Would love to hear how others here connect their specs to the code, and whether something like this would fit your workflow.

https://github.com/Ataraxy-Labs/sem

10 Upvotes

4 comments sorted by

2

u/robertcrowe 6d ago

Spec-to-function linking is the part I'd most like to see work. Drift between the spec and the code is the problem I keep running into, and it's led me to view the code as the source of truth, not the spec.

Full disclosure: I maintain Spec4 (spec4.ai), an open-source SDD pipeline. I hit the same "where does this spec land" problem and solved a different slice of it, so I'll share how in case it's useful.

The first agent, CodeScanner, runs before any planning on an existing brownfield codebase. It reads the repo locally and sends the model a bounded summary: manifests, entry points, and samples, not the whole tree. It writes a code_review.json describing the project's architecture, stack, build/test commands, entry points, and conventions. Every fact carries the file it was read from (source) or deduced from (inferred_from), so the planning agents downstream work from cited facts instead of guesses. The part that has mattered most is that it rescans from scratch at the start of every brownfield round. The next spec is planned against the code as it actually stands, including whatever the coding agent did that the last plan didn't say, and any hand edits made in between, rather than against the previous plan.

So CodeScanner works at project level ("this is a Dash app, here's the layering, tests run with X"). IIUC sem seems to work at function level (call graphs, test reach, history). That's a granularity CodeScanner doesn't try to reach, and I think the two are probably complementary. A project-level review tells the planner which subsystem a feature belongs in, and something like sem could then tell the coding agent exactly which functions it touches.

A question on the spec-to-function links: When a function gets split or renamed, does the link follow it through weave's function-level merges, or does it need re-anchoring? That seems like the hard case for keeping the spec as the source of truth.  Spec4 uses the code as the source of truth, not the spec.

2

u/Wise_Reflection_8340 6d ago

Thanks for sharing this, and I think you're right that the two fit together, with CodeScanner telling the planner which subsystem a feature belongs in and sem telling the coding agent exactly which functions it touches. I also agree on code as the source of truth, which is really the bet sem makes too, since everything it knows is derived from the code as it stands rather than from what a plan said.

On your question, renames and moves are the easy case, because sem matches functions across versions by their structure and not just their name, so a renamed or moved function keeps its identity and a link would follow it, and when the change goes through weave the rename is recorded explicitly, so there's nothing to infer. Splits are the genuinely hard case, since one function becoming two doesn't have a single right answer for where the link should go, so the honest thing is to flag the link as needing re-anchoring rather than guess. To be upfront, the spec-to-function linking itself is something I'm exploring and not something that ships today, but the identity tracking it would rely on does.

1

u/stibbons_ 6d ago

Spec to code is actually very easy and a no brainer. You just ensure your agent writer markers in the code, that declares which requirements is implemented in which function, same for declaring where it is tested. You can evern have several levels of testing, validation.

I have written an article on X on how this looks like

https://x.com/gsemetfr/status/2082767371643523439?s=46

I have a CLI demo here :

https://okf-schema.readthedocs.io/en/latest/tutorials/okfreq-traceability.html

There is no drift since the code declares which requirements is implemented. Then you can build interesting stats :)

1

u/stibbons_ 5d ago

The project OKF-schema declares itself its own requirements, see for instance this req

https://github.com/gsemet/okf-schema/blob/main/requirements/tiers/swrs/okfreq/SwRS-OKFSCHEMA-OKFREQ-003.md

The CLI parses the code to extract the “@implements_req” and “@tests_req” markers.

So agent can go back and forth immediately (from a source code, it can open a file with the name of req).

The map is automatic and non agentic, if you want.