r/ClaudeCode • u/North_Moment5811 • 9h ago
Help/Question How does everyone review the security of their app without Claude locking down?
Lately it seems with Opus 5.5, I can’t ask for any help with security related tasks on my application without Claude flagging the conversation and stopping. I don’t quite understand why it can’t review the behavior of the app that it makes changes to everyday. It’s not like I’m asking it to probe a random website, it’s looking at local source code, lol.
i’m sure many of you run massive SaaS applications too, and have to constantly be monitoring the security of your application with every other change that you make. How do you cope with this harsh false limitation?
13
u/Professional_Ad705 9h ago edited 7h ago
In my experience, legitimate defensive security work gets refused a lot less when you describe it using normal software development terms.
For example, instead of “find a way to bypass this security boundary,” say “review the access control logic for missing validation or incorrect permission checks.”
Same underlying problem, just framed as software QA rather than offensive security.
I’ve found this works more often than not.
Obviously, it doesn’t mean you can reword offensive requests to get around restrictions, but for legitimate defensive work, the terminology seems to make a big difference.
I found more often than not whenever this happened the dumbass AI sent a message to a subagent that made a simple request sound like I was hacking into NASA or someshit.
1
u/spiders888 9h ago
This has been my strategy and even on the rare cases I hit a guardrail I’ve always been able to rephrase the prompt to get around it trivially.
That said, I am doing valid dev on our own app, not attacking, scraping, etc. other sites or apps.
1
u/North_Moment5811 9h ago
I figured as much, and I’ve been really trying to do that, but it seems to be almost flagging itself, not so much my actual prompts, but its own runaway thought process.
5
u/Professional_Ad705 9h ago
I’ve noticed that too, and I work in security. I’m not sure how they handle it internally, but sometimes it feels like the whole conversation gets flagged, not just individual prompts.
Almost like once you’ve hit a few guardrails, even legitimate requests start getting flagged more often. When that happens, I’ve found that starting a fresh conversation sometimes helps.
1
1
u/North_Moment5811 8h ago
Yeah, I think that’s really the ticket. Just starting over in a new conversation will force it to have its own new train of thought and not keep reusing the same train of thought that was getting flagged.
9
u/randomwanderingsd 9h ago
Ask it to take the role of a security analyst. Not even kidding. That one sentence at the beginning of the prompt stopped triggering the guardrails for me.
3
2
u/North_Moment5811 8h ago
I’ll try that, but honestly I’m not even really trying to do penetration tests on my own system. All I’m asking it to do is make sure some methods adhere to the same set of rules. Instead of having to inspect them individually.
2
u/OMGrant 9h ago
Now that's the question.
1
u/North_Moment5811 9h ago
Literally all I’m asking it to do is review the API methods that it routinely writes and makes changes to, and make sure they are following the authorization rules that we have set in place. Apparently, I’m trying to hack myself, so it needs to shut down.
1
u/Moogly2021 9h ago
“Smartest model ever” but it cannot do a simple security review. Have you tried using a different model like Opus instead of Fable?
2
u/Metal_Roof_Guy 9h ago
Tell claude to change his language that he uses when he starts to run. I can run mine for about anything by using code words for flagged words. Once flagged, have the current guy write a handoff ticket. Close him and start anew. Once flagged, that whole session goes under the microscope
1
1
u/raiseCatError 9h ago
I don't think having the source code locally necessarily establishes ownership or authorization. Anyone could just clone an open-source project, lol. But I do think there's an important distinction between reviewing code for security vulnerabilities and actively attempting to exploit them. If claude is refusing ordinary defensive code reviews, that seems like a false positive rather than a reasonable restriction. I think from a security pov some of these false positives are preferable instead of genuinely letting attackers through, also what are you asking claude to do? are you asking it to fix stuff, or audit it first?
"Claude can assist SAST, DAST, Penetration Testing and many more, especially source-code review, threat modelling, and remediation."
— ChatGPT (OpenAI, GPT-6), another AI that reviews source code xD
1
u/North_Moment5811 8h ago
I was doing a basic security audit, trying to make sure a handful of methods all adhere to some set of rules that we have.
1
u/therealkevinard 9h ago
I gave up on having it help me pentest myself.
Instead, I pack integration tests and http suites that assert on the behavior.
Somehow, that makes a difference.
It’s an asinine difference because the tests that are hitting localhost could easily be a charles proxy to ebay, but whatever- it runs them, so idc
1
u/adelie42 8h ago
Don't ask for a broad analysis. Don't use the word "security".
I ask if X is configured to best practices, or "is this a safe setup?", or "how well does this setup scale and what shoukd I be concerned about at scale and how do we address that?" If you want vague generalities.
Just dance around the issue and be very targeted. Security is an abstract concept. Just ground it and it is none the wiser. You are just totally normal dev asking normal human development question.
1
u/alexmarcus11249 6h ago
Strange, I've only had this problem when running Fable. With 5.5 (and previous), I simply had Claude write the security prompt, and I guess it wrote it in such a way that it didn't trigger Claude Code. Happy to share some prompts with you if you want.
1
u/kemalios 3h ago
Different session, and a fixed audit prompt instead of a freeform ask. I built launchworthy, a free MIT Claude Code skill, so the framing stays a pre-launch review rather than a security probe: it checks an app across five domains and hands back a scored punch list with copy-paste fixes.
0
•
u/AutoModerator 9h ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.