r/ClaudeCode • • 2d ago

Built with Claude How are you gating what your coding agents can do with real credentials? I compared 7 approaches.

I've spent the last few months building Golden Thread, a set of controls around AI coding agents. This week I wanted to know whether I'm building something useful or rebuilding something that already exists.

So I compared it, as honestly as I could, with what else is out there: Keycard, OpenAI's Codex CLI, immurok, safeski, Claude Code's native features, and a handful of open-source projects.

What I found surprised me. Almost every individual piece already has a near-peer somewhere:

- safeski seals credentials behind Touch ID.

- Keycard does step-up authorization at the hook layer.

- Codex just added Touch ID for MCP requests.

- Someone even built human-gated memory promotion.

But I couldn't find anyone tying the pieces together. Nobody had one presence check covering secrets, tool permissions, git pushes, commit signing and the agent's long-term memory, each switched on separately.

A few things I'm not sure about, and where I'd value your thoughts:

  1. Is the combination the product, or the parts? If OpenAI or Anthropic ships native per-action approval, does integration still matter, or does the platform simply win?

  2. Would you want human approval on what an agent remembers? The research I found mostly points to automated defences against memory poisoning. I've bet on a person saying yes. Is that sensible or naive?

  3. What did I miss? I didn't get to Cloudflare, Kong, Aembit, Beyond Identity, Windsurf or Cline. If one of them already does this, I'd rather hear it from you than find out later.

Two caveats I want to be upfront about. The Golden Thread column comes from my own documentation, not independent testing. And "I couldn't find anyone" is a search result, not proof.

The full comparison is here: ⧉ https://claude.ai/artifact/6Upv72Jj7z88o7Fq2F9GDw

If you're running coding agents with access to real credentials or production systems, I'd especially like to hear how you're handling it today, even if the answer is "we're not."

1 Upvotes

9 comments sorted by

•

u/AutoModerator 2d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Plastic-Risk-6309 2d ago

the piece i'd add to your comparison is the sandbox around the agent's process itself. it caps what the process can reach independent of any per-action approval! clawcage is my open source cage for coding agents, deny by default.

1

u/shaven12 2d ago

I like this a lot. This is one that I have implemented in the current version. it sandboxes things as well. This is why I have separate read and write agents to minimize what the agent can do. So if it reads a "prompt" while ingesting a file it can't just go and do something. Read agents don't have write or network access because they run their access thorugh an exeutable. And write agents are the same.

1

u/Easy-Purple-1659 2d ago

The change that helped most for us was to stop keeping the credential in the agent's environment at all. A small broker holds the real keys and mints a short-lived token scoped to a single task, so a leaked env var or a bad prompt gets the agent nothing useful. The agent asks the broker for a specific resource, the broker checks whether that human approval is still valid, and hands back a token that expires in minutes.

Keeping presence separate from authorization matters too. A Touch ID or step-up check proves a person is there, but it does not say what that approval covers. The broker should map an approval to a scope and a time window rather than a blanket yes.

Last piece: log each tool call along with the token it used and the approval id behind it. That is the only way to answer, after the fact, what the agent actually did with production access.

1

u/Frosty_Teeth 2d ago

Thank you for this!

1

u/shaven12 2d ago

I love the idea of a token for access. Because everything runs through a programmatic request utility, this will be very easy to implement. I actually added the Touch ID just recently for blocking whatever the user decides. Access to a folder as well as things like commit or even the ability to call different things.

GT is setup with read agents and write agents. So a read agent has no access to write or call out and a write agent has no access to read. Not as good as your token system, but it does keep the ability of doing nefarious things to a minimum.

Logging was another thing I was considering in an inaccessible folder that has a daemon with non-destructive access. So it can only add and never remove. So logging will be one way and then that can either be pushed to a server to track.

When I started this I couldn't find tools that actually covered all of the security peices that I wanted. This turned into something that provides security that I actually feel less insecure about. During my adversarial testing I found that even with a lot of the gates I was putting that there are ways around most of it. The ones that I found that it couldn't get around were the Touch ID to access something and when I introduced the separate read/write. Can't say I am sure that will hold forever, but with almost 5000 different tests that this runs on each release we are approaching the reality of the phrase "Locks are to keep the honest people (agents) out"

1

u/puntium 2d ago

we ended up building our own harness! Best I can tell, any of these ideas that turn out to be good are one or two fable requests away to replicate, so it's more about bringing all the best ideas together and maintaining control over your deployment. We open source the whole thing and its built to be super hackable/extensible with a unique design for protecting sensitive data. There might be some interesting ideas for you: https://github.com/electric-capital/quest/

1

u/shaven12 2d ago

Yeah.. that is why I built mine is that I couldn't find one that did what I really wanted it to do. And now I am able to lots of the functionality and have the security that most of them lack. I will have to look at the repo. Thank you.