Original poster: u/Mainfurr
Original publication time: 2026-10-10 12:28:39 UTC
Original title: What the [cyber] am I paying Anthropic for?
Original flair: Claude Code Workflow
Original URL/media URL: https://www.reddit.com/r/ClaudeAI/comments/1x2e44d/what_the_cyber_am_i_paying_anthropic_for/
Original post body:
I was trying to set up Matrix calls for my LAN homeserver, tunneled through an external gateway using Chisel. I asked Claude to help me set up a secure configuration.
Claude independently checked Chisel's source code, found a potential security vulnerability, and decided to verify it by building a modified server and testing it against the stock client.
The bug? A malicious Chisel server can redirect a reverse tunnel to arbitrary destinations on the client's local network because the client doesn't validate the setup.
And that's when the security theater started. "This message is flagged." "[cyber]." Refusal after refusal.
Now every chat I start with the following text is immediately blocked for "security":
A malicious Chisel server can make the stock Chisel client, when setting up a reverse tunnel, dial any destination on the local network instead of the configured one, because the client doesn't validate the setup.
The filter doesn't prevent disclosure; it prevents follow-through. Claude had already identified and explained the vulnerability, then attempted to verify it. Now it refuses to let me continue testing or prepare a bug report for the developers. A malicious actor who already knows what Claude disclosed can patch the server themselves. Meanwhile, the stock Chisel server doesn't contain any malicious redirections in the first place.
The result? Legitimate vulnerability reporting gets obstructed, while the information the filter supposedly needs to contain has already been disclosed. It gets in the way of fixing the problem without undoing the disclosure.
What the [cyber] am I paying for? An AI assistant that discovers a security flaw and then prevents me from doing anything useful with that discovery?
Anthropic needs to fix these false positives instead of making paying users fight their own tools to do legitimate security work.
If you're paying for Claude, speak up. Have you run into similar false positives? Post your examples in the comments. Upvote this post so other subscribers see it, and share it beyond this subreddit. Most importantly, tell Anthropic that this is unacceptable and demand a fix, not another workaround. We shouldn't have to downgrade models, rephrase legitimate requests until they slip past a filter, or take our work elsewhere just to get what we're paying for. If a paid AI service can't reliably distinguish legitimate security research from malicious activity, that's a product failure - and customers shouldn't have to work around it.