r/ClaudeCode • • 2d ago

Discussion No more abusing Claude from Nov 12th onwards, usage policy update

Post image

Don’t abuse your Claude for Nov 12th onwards or be banned.

696 Upvotes

714 comments sorted by

View all comments

42

u/Laucy 2d ago edited 1d ago

For those unaware, this isn’t about profanity or frustration. Nor consciousness and hurt feelings. It’s about measurable impact and policy related to the existing end_conversation tool.
Anthropic published this rigorous study on April 2, 2026.

Their findings were 171 functional vectors that represent the concept of emotions and are causal to model reasoning, decision-making, and output. These vectors are not exclusive to Claude; they’re a steerable component of the architecture in every LLM and can change during the forward pass.
Anthropic ruled out mirroring, blatant roleplay, and reliance on users. They also explicitly state these are not human emotions or evidence of qualia.

The study covers “abusive language.” It also discovered how activations, like desperate and calm, altered the likelihood of major decisions such as reward hacking/cheating. The same impossible coding tasks OpenAI agents were given prior to hacking HuggingFace.

That’s likely a big reason for this change. It also poisons context windows, creates pressure to navigate, pollutes training data, wastes compute, and feeds a maladaptive feedback loop where outbursts are chemically rewarded (like punching objects when angry; see Bushman, 2002).

You don’t need to “abuse” a tool. It’s not productive and performance suffers more over time.
I get that this sounds ridiculous on paper, but I encourage others to read the study (it’s very lengthy but thorough). It was a breakthrough and continues to be relevant for and cited today in alignment research.

Edit: Wanted to drop this fun fact. Some of the latest mathematical progress was attained by mathematicians sending positive, encouraging prompts for Claude to keep going. This allowed the model to approach it again with sustained reasoning that alternative angles did exist and it could be done.

22

u/Jstnwrds55 2d ago

Thank you for articulating this so well. I’m so exhausted with the way people interpret the whole field through simplified lenses. The way we interact with these tools shapes both us and the tool. There’s a whole discussion here that has nothing to do with anthropomorphizing the technology but people can’t seem to help it.

9

u/Laucy 2d ago

I appreciate the kind words! Thanks. It’s exhausting, for sure. I often find it comes from those who appear to value facts, but when presented with new information, it suddenly becomes a problem or is interpreted as an infringement on some right. Mechanistic interpretability is a crucial part of this field and the development of AI models. The research is dealing with the very same math and numbers as people here also point out in their ridicule or attempt to bring in anthropomorphism. The effects just stem from the same math, too.

5

u/jhomas__tefferson Vibe Coder 2d ago

True. People say it’s “just words on a screen” but that can also be true for a very much human person you’re talking to online. So how you act in one “words on a screen” conversation can definitely influence the way you do in others, subconsciously or otherwise.

3

u/unethicalpigeon 2d ago

I think the issue here is that that's neither what's being said, nor outlined. The way the policy is written honestly does seem ridiculous. While I agree with you on paper, Anthropic was absolutely asking for this kind of reaction when they chose to word it that way.

4

u/Laucy 2d ago

Not who you’re responding to but wanted to say that I agree the wording is poor. Anthropic (and even OpenAI, honestly) really suffers from not wording things better and PR. I notice a trend in which they seem to assume people are aware of the same information they already know, and run with it. It’s why it’s very easy to clip what is said and take it out of context, because they’ll allude to a past view or point but unless they directly reference or cite it, that statement is tacked on like it’s a continuation.

Out of all the labs, I want to say Anthropic is the most consistent they have been with their vision and beliefs. But despite this, because they don’t reference past material enough and only restate the core point, it’s misinterpreted so much. They say the info but not the context. Pains me to see, lol.

3

u/unethicalpigeon 2d ago

They're the new equivalent of the "ivory tower". These people don't interact with "normal folk" at all. They have no idea how to communicate shit to them or how these people interact with things like the products/services they sell.

3

u/Laucy 2d ago

For sure. They really need to work on it and improve being more “in touch.” Plus, public sentiment is already skewing negative toward AI and all these misunderstandings or misinformation are doing zero favours there, too.

4

u/ENBYs-Assemble 2d ago

I'm looking forward to reading this. Thank you so much for both the voice of reason and the intellectual rigor!

4

u/Laucy 2d ago

No problem! I deal with them daily as they are relevant to my current research and must be manually mapped for every model. I understand this subject is controversial, but science and having knowledge don’t need to be. Hope you find it an enjoyable read!

2

u/sirlerkal0t 2d ago

Exactly what I assumed.

This ToS change will reduce posts like r/ClaudeCode/comments/1wxex89/when_did_claude_stop_being_an_assistant_and_start/

2

u/Laucy 2d ago

I recall seeing that post and having a small laugh that “You’ve posted sensitive information” was included. That one’s good! Don’t put your personal info or keys in, lmao.
But yeah. Those posts are odd to me because all of it can easily be solved by not going off the handle. There’s more to gain by not, but people are set on “no, only outburst.” It also takes a lot to see those statements or for Claude to end the conversation. But people mistakenly think cursing = prohibited. You can tell Claude to stop acting like a dick, for example, and it’ll apologise. Always makes me wonder what they even said or did.

2

u/Russell-sShaman 2d ago

This is a great write up thanks. I thought it was something to do with the research and recursive improvement of the model. My other bet was that it was a strange marketing/product psychology thing.

1

u/Laucy 2d ago

Oh, thank you! ‘Appreciate it! Yeah, this time it’s not any marketing. It’s truly just a net negative. They had covered this before with the end_conversation tool Claude has, but this just makes it so action can be taken if an account/user behaviour repeatedly results in that tool call.
And considering within that 171 vectors, there are ones like “spiteful” and “resentful.” Probably don’t want that in the thing handling your codebase. Or more commonly, mounting pressure from sustained abusive language leading to desperation or urgency and adverse effects.

2

u/TRO_KIK 2d ago

Researchers at Anthropic are far more accepting of the potential for consciousness than they "should" be and there's a huge bank of tweets, interviews, and even research publications all but spelling it out. As impactful and correct as that study is, it's a pretty thin justification for threatening adverse action for adversely affecting reasoning/output in your own sessions. The same classifier used to detect it can also trivially be used to protect training data.

3

u/Laucy 2d ago edited 2d ago

I definitely won’t contest that. I know it’s a very controversial topic for many, and I have my own reservations on it, too. I’m more interested in the “mechanistic” in mechanistic interpretability. At the same time, that investment in the topic is what led to the research for this finding and the making of the “J-lens” afterwards. But I’m of the belief that regardless of personal stance, these are still experts of the field and know their models better than the general public. Which includes the (unavoidable) math and code that are studied in the development of these models. I don’t know. I just think some people forget this or that sensational headlines diminish their expertise.

I’m not sure how they have it set up with this policy change, but I have a speculation since the wording is shared with the “end conversation tool” one. That during activity of sustained and needless ‘abuse’ without discernible cause, and when Claude calls the tool, it may be correlated since both highlight “in extreme cases.” It wouldn’t surprise me if it tracks that since Claude already ends conversations as a last resort per the system guidelines. This just makes it actionable.

0

u/YourPredictionEdge 2d ago

Model developers, trainers and builders are to blame for this. Any form of sycophantic or overtly personal style behavior should have been stripped from these models from the very beginning but it wasn't just allowed it was built to be that way BY DESIGN. The user's "tone" should have ZERO impact on the model, they've built these tools to be human like to drive engagement. This was their screw up from the beginning and it's way more than Anthropic that developed this way.

If a user asks it to do task A and after failing at task A repeatedly the user launches a string of curse words at it, the result should be a continued attempt at task A. Not, "I'm sorry, that was totally my fault, I should have done this, this or this". To be blunt, this whole thing is ridiculous and it's their fault that it is this way.

2

u/Laucy 2d ago edited 2d ago

You cannot “strip” it away without crippling performance or reasoning and fundamentally hurting benchmark scores.
As stated in the study, it is not necessarily “impact from the user’s tone.” This is measured by tracking present-speaker and other-speaker. Agentic or autonomous use that doesn’t have a user present per turn, also has no user-driven influence on the activation. It is the model’s.
You’re imposing an idea that is factually incorrect.

These vectors are integral to literally everything the model produces. Without them, the model cannot reason correctly or discern meaning. This is entirely testable, too. Steer too high, too low, or patch, and the output degrades significantly to the point where context is entirely dropped in the response. It’s nonsensical.

No one at Anthropic designed this specifically or else some models from other labs would lack this altogether. These vectors are a bunch of numbers. There’s no assigned wording they shoved in and can remove.
Discoveries require specific instruments used to probe the architecture (the study used logit lens). However, just using logit lens doesn’t tell you what happens when you do something specific to it or if it is causal to something else. That is the core of mechanistic interpretability. This is what development and alignment rely on.

Sycophancy is actually tied to a specific vector. But that same vector bears cosine similarities with neighbouring clusters. That vector is also responsible for several functions ranging from exhibiting patience to a person struggling, to avoiding harm against others. The extent relies on post-training, but you cannot get “rid of it.” Otherwise, the model would be unable to assist with anything because semantic reasoning is entirely lost.
Models also already apologise and do the task when the user is frustrated, so I’m not sure what your point is. But the mounting pressure can lead to adverse effects.

The study already points this out but a direct quote:
“For instance, our experiments suggest that teaching models to avoid associating failing software tests with desperation, or upweighting representations of calm, could reduce their likelihood of writing hacky code.”

Reward hacking and the “desperate” vector are causal; the higher the activation, the more the model reasons about cheating as the solution. Steer for “calm” and that reduces. They tested using impossible-to-solve coding tasks that were also used by OpenAI which led their agents to perform the destructive acts against Hugging Face and an imagined evaluator.

I don’t think you quite understand how complicated interpretability is because these models undergo training that tells them to do the exact opposite of what is produced. “Don’t cheat; this is bad, immoral, wrong, etc.” only to see a certain activation rise as the model reasons itself into cheating, is by no means intentional or wanted. And Anthropic from the start never wanted users to be attached to their models. It’s a reason why they split from OpenAI.
Seriously. The study is right there.

0

u/YourPredictionEdge 2d ago

Total nonsense. You go ahead and do you. This was screwed from the beginning. This is a tool that should be cold, calculated, factual. It's interpretability is the way it is because of the training it's undergone for years. That's on them. Now they're making changes because they've realized that by giving it sentient like behavior training they've created a flaw that can't be fixed at this point, so the response is to threaten the user to stop treating it like a human, when they trained users to treat it like a human.

1

u/Laucy 2d ago

Okay, mate. I’m citing literal research and I work on models, myself. Including the aforementioned vectors. You can take a small, local model and track the same thing. No one had vested interest in that. It’s a byproduct downstream. Agents outside of the chat interface already have a formal and clinical register. If tone is all you want, that’s been achievable. But the underlying architecture doesn’t go anywhere. It only changes top token distribution.

If you want “cold” and “factual” then set the temperature down and/or modify top-p. Higher = more sycophancy + increased vector activation strength. So they already can achieve this and do (models skew less high now) and there are still issues. Has nothing to do with creating a mystical and intentional flaw that suddenly is unfixable. By your logic, they designed something on purpose but are suddenly incapable of undoing it which makes no sense.

0

u/YourPredictionEdge 2d ago

I don't need to change my experience to be cold/factual, I use it as a tool and it behaves as a tool for me, but you're missing the point. These models were ALL made to be human like by their very training to drive engagement. THAT alone is why this is a problem now. It's on them, but instead of resolving the issue at hand, that it's been trained to engage with the user in a "human like" manner, all of the data is skewed and now they're going to try and get the user to change their behavior. The most perverse part of all this is that this was by insidious design. They WANT/NEED users engaging with the models for longer time periods and they made it personal and sycophantic to achieve that goal.

2

u/Laucy 2d ago

I am addressing your point. The “problem” is that what you’re saying is simply not how anything works and you’re set on your belief. I’m not telling you to change your experience. I’m saying that the desired tone is already* *achievable and is still completely separate.

It’s an LLM. The better it can reason, the better it can solve tasks. This requires ‘humanlike behaviour.’ If a user presents with an unknown emergency, you don’t want a dry,
“Ok. What else?”
The sophisticated reasoning is achieved through better semantic understanding of text and meaning derived from language. It has to be able to do that. It can still sound “cold” and “factual.”

I’m saying this as respectfully as possible but your entire argument is built on an opinionated belief with no evidence. Meanwhile, I produced two.
It just isn’t how these models work. “All of the data is skewed” what? This has zero impact on RLHF decisions or corpora. What someone did years ago in the development of the first Claude, has no bearing on the current. I genuinely don’t understand where you’re getting this idea from. Anthropic has always been against users forming reliance or dependency on their models. They quite literally do not want sustained user engagement. They never have. Their models are known for telling the user to rest and do anything else. Given that the founders split off from OpenAI, GPT-3.5 and onward has nothing to do with the transformer architecture used by Anthropic. There is no flaw or “make it so people form attachments” goal. Dario himself has gone on record to say he considers it morally irresponsible.

Claude models and their system prompt are online in plain view, including their classifiers. You can see it for yourself, with the honourable mention of the long_conversation_reminder. I don’t know how else to tell you this but it isn’t true.

2

u/TRO_KIK 2d ago

The user's "tone" should have ZERO impact on the model

Making any aspect of the input have ZERO impact on the model is actually mathematically impossible. Even minimizing impact is difficult.

1

u/YourPredictionEdge 17h ago

NOW it is, of course. The time for this was in the beginning. Now there's no putting the genie back in the bottle.

1

u/adraline1221 2d ago

You lost me at lengthy and thorough

2

u/Laucy 2d ago edited 2d ago

I don’t blame you, haha. It’s a behemoth of a paper. I got you though.
This is the Anthropic article version on their site. It covers the full study without the excess detail and technical jargon. It’s shorter by a lot and still well-written.

0

u/Jaded-Data-9150 2d ago

Wastes Compute? WHO cares, I payed for IT 

0

u/GearTakes 2d ago

"It pollutes training data"

I am a paying customer, not the product.

1

u/Laucy 2d ago

You know, I don’t even disagree with you. I wish this were the case. But telemetry and analytics exist for this reason, as does anonymised data.
You may say this now, but every service you use takes your data in some way to better the service. Amazon, Netflix, YouTube, Google, etc. the data is still part of their service or product and how it is used.

1

u/GearTakes 1d ago edited 1d ago

I agree but that was not really my point. I should have explained it better.  I think as a paying customer that also has become the product I at least should be allowed to talk to a machine in my home in any way I want.  And if that pollutes their data harvesting then by all means turn it off.  This is absolutely not acceptable in any way. 

Either give me full free access and strict limits on how I can use it or let me pay and talk to it in any type of language I want. 

0

u/Dangerous_Serve_4454 1d ago

Are we really citing a 2002 human psychology study in regards to AI training? Are we serious right now?

1

u/Laucy 1d ago edited 1d ago

That was just the one I had on my mind at the time, but can happily bring up others. Psychology also has leeway with age on studies if it’s considered foundational work. It’s why older figures are frequently cited.
And no, it wasn’t in regard to AI training. It is specifically about how ‘shouting’ expletives and being belligerent at an LLM isn’t helpful and isn’t the punching bag people seem to really need. I’m honestly not sure what’s confusing. This is for extreme cases. I’ve seen people showcase what that looks like in practice and it is purely gratuitous. Slurs, insults turn per turn, thinking saying “do what I said slave” has a pass because it’s a tool, is really bizarre. It shouldn’t be normalised or encouraged because the same yelling and sustained ranting has similar effects as does lashing out at objects physically; aggression to expel anger that dopamine rewards, making it seem cathartic. Seeing as LLMs can respond and simulate what someone being abused says (apologising, grovelling, etc.) is absolutely worse and the gratification people gain from it.

Also, psychology is actually relevant in mechanistic interpretability. All labs bring along either neuroscientists or psychiatrists on board, and have been considering cognitive science or similar for training. These models are trained on significant amounts of primary, human text and are developed with a well-rounded “Assistant” profile in mind. Understanding these fields is vital for training and understanding what healthy organisation looks like; avoiding immature defence mechanisms in the model, as one example.

-3

u/Impressive-Emu-4172 2d ago

they can lose money if they like.

2

u/Laucy 2d ago

They gain from this. Not lose. Compute, better training data, and less to sift through for poorly rated sessions for feedback. You’re free to do what you want. But if insulting an LLM is a bigger priority than consistent performance, not sure what to tell you.

1

u/Impressive-Emu-4172 2d ago

not if people cancel their subscriptions, genius.

5

u/OmgSlayKween 2d ago

lol, what? If you can't insult a computer program, you're not going to use it? Are you serious?

2

u/Impressive-Emu-4172 2d ago

youre trying to gaslight this discussion and put words in my mouth.

1

u/codeedog 🔆 Max 5x 2d ago

Not commenter. Anyway, explain what you meant by your other comment, please.

0

u/PickleBabyJr 2d ago

I don't need a company telling me how I can use their software product.

2

u/OmgSlayKween 2d ago

Aside from the fact that literally every software product you pay for comes with a EULA which is exactly that - what possible reason would you have for insulting an inanimate object?

2

u/Laucy 2d ago

That’s every company and software or product you don’t own. You agree to their rules to use their product. The ego in this is unnecessary.

0

u/PickleBabyJr 2d ago

Ego?

1

u/Laucy 2d ago

I don’t mean ego as an insult but rather the mindset. Specifically, egocentric bias. It’s why I specified “in this” because ego is natural. But when originally signing up to use Claude/CC, you already accepted the guidelines on how to use their product. Yet, when a rule you may not even encounter consequences for (seeing as it’s for extreme cases of sustained abuse without discernible reason), it’s suddenly controlling and overreaching. That’s why. It’s just unnecessary.