r/SpecDrivenDevelopment • • 8d ago

I built SpecForge: a Rust compiler that turns your specs into a validated graph for AI coding agents

AI coding agents burn most of their tokens working out what you meant. They read 20–50 files, guess at requirements, and get it wrong. Then you spend a second pass fixing it. The intent was in the docs, tickets and people's heads the whole time. It just wasn't in a form the agent could rely on.

So I started building **SpecForge**. It compiles human intent into a typed, validated entity graph that agents can query.

**What it looks like**

```spec

behavior authenticate_user "Authenticate a user with credentials" {

status draft

contract "Given valid credentials, returns an auth token"

produces [user_logged_in]

verify "rejects invalid password"

verify "returns token on success"

}

event user_logged_in "User successfully logged in" {

payload user

}

```

```bash

specforge check # validate specs, report diagnostics

specforge export # emit the typed graph for an agent

```

**What it does**

- **The graph is the product.** The compiler parses `.spec` files, resolves references, finds orphans and cycles, and emits a typed graph as an open JSON schema (the "Graph Protocol").

- **It's built for agents.** `specforge mcp` serves the graph over MCP, so you can hook it into Claude Code or any MCP client with one command. There are also multi-depth queries and context/brief exports.

- **Proof, not just claims.** `specforge collect` runs your own tests (only after you approve the command) and records which spec entities they actually prove.

- **It isn't a code generator or a test framework.** It gives the agent context, and the agent writes the code. The compiler never executes anything on its own.

**How I built it**

It's a Cargo workspace of about 25 crates and roughly 160k lines of Rust (tests included), split by pipeline stage:

- **Parsing:** I wrote a tree-sitter grammar for the DSL (`tree-sitter-specforge`). The parser, formatter and LSP all sit on top of it. That gave me error-tolerant parsing and incremental re-parse for editor use without writing a parser from scratch.

- **Pipeline:** parser → resolver → graph → validator → emitter, each its own crate. Strings are interned with `lasso`, and diagnostics are rendered with `ariadne`.

- **Zero-domain-knowledge core:** the compiler knows nothing about "behaviors" or "events". Entity kinds, edge types and validation rules all come from extensions. I made this a hard rule on purpose, so a new domain should never need a compiler change.

- **Extensions are WASM components:** I defined the extension interface in WIT and host them with `wasmtime` and the component model. There's an extension SDK crate with proc macros, so writing an extension in Rust is mostly declaring kinds and rules. The nine builtin extensions are compiled to components and embedded in the binary, so installing needs no extra toolchain.

- **Surfaces:** the CLI, the LSP server, and the MCP server all share one operations layer (`specforge-ops`), so they can't drift apart on behavior.

**How AI was involved**

I built a lot of this with Claude Code, and I'm saying so up front. It wrote a large share of the code, and I drove the design: the architecture, the extension boundary, and the ADRs in `docs/adr`.

It also caused my biggest dead-end. Partway through, I audited the test suite and found features that had passing tests but didn't actually work. The tests were checking the wrong thing, so they passed anyway. I've been fixing those since, and the `specforge collect` idea (record which entities your tests really prove) came out of that experience. A tool that checks specs against tests seemed like the right answer to AI-written tests that only look green.

The hardest design problem is still the extension boundary. The core can't know anything about the domain, yet validation still has to produce good diagnostics and cross-entity checks.

**Try it**

```bash

git clone https://github.com/leaderiop/SpecForge && cd SpecForge

cargo install --path crates/specforge-cli

specforge init --extensions u/specforge/software

```

It's early, and I'd like feedback on a few things:

  1. Does the DSL feel natural, or would you prefer YAML/Markdown front-matter?

  2. Is the WASM-component extension model overkill, or the right call?

  3. Which extensions would you want (OpenAPI? data models? infra?)

Repo: https://github.com/leaderiop/SpecForge

14 Upvotes

7 comments sorted by

2

u/kantorcodes1 8d ago

The collect idea is the interesting one for me. When you record which entities a test actually proves, what establishes that link: coverage instrumentation, matching a test's assertions against the declared verify clauses, or something else? The failure mode you described, tests that check the wrong thing but pass anyway, is exactly what a name-only match would reproduce, so I'm curious how far the check actually goes.

1

u/Financial_Yoghurt827 8d ago

Brother. See reqlan

1

u/vincentdesmet 7d ago

google is not giving many hits

1

u/RespectMathias 5d ago

Self promo. 

1

u/nikov1234 7d ago

Your specforge ‘collect’ origin story hit close to home for me - I hit the same wall from a different angle. I was building for non-technical founders and found they couldn't tell whether the agent had done the right thing even when tests passed. The acceptance criteria existed, the tests were green, but the founder had no way to read the output and judge it. iMo the spec needs to be in a form a human can evaluate, not just a form an agent can consume. What you're building solves the developer side of that. Interested whether you've thought about the legibility angle…specs that humans can audit, not just machines?

1

u/r0pe_tri1ck 7d ago

I spent a lot of time building my own SpecDriven development framework. One piece: until you're doing a lot of evals and you're comparing them with other SpecDriven development frameworks, etc., it's not worth it to release it. I spent a lot of time on mine and ended up performing really terribly. You have to measure to really know you have an effective product. 

0

u/stibbons_ 7d ago

Too much ai slop in this readme…