Anthropic has updated the terms of use for its services and prohibited users from engaging in “systematic and unjustified abusive or demeaning behavior” toward Claude AI models. The rules take effect on November 12, 2026. Violators risk not only having the bot end the conversation, but also having their account permanently suspended. The company attributes this move to research into “model welfare” and efforts to identify signs of consciousness in neural networks.
What exactly has changed in the rules?
Anthropic has officially introduced a ban on “persistent and excessively abusive or demeaning behavior” toward its Claude language models. The new rules apply to absolutely all users—from subscribers to the basic Claude.ai and Claude Code services to enterprise customers and third-party developers using the API.
A tiered system of penalties is in place: the chatbot itself may end a conversation early, while Anthropic's security team may issue warnings, restrict access, or permanently ban accounts. Users are also explicitly prohibited from circumventing bans by creating new accounts.
At the same time, Anthropic emphasizes that these penalties apply only to extreme cases and systematic harassment with no clear purpose. Ordinary human frustration over coding errors, criticism, research, simulations of complex scenarios, or “dark creative themes” do not fall under this ban.
How does the neural network protect itself?
The conversation termination mechanism began to be rolled out with the release of the Claude Opus 4 and 4.1 models. If a user is persistently aggressive or demands that the ethical guidelines be violated, the model tries several times to redirect the conversation before simply ending the dialogue. The thread then switches to read-only mode. The only exception is when the user indicates a risk of harm to themselves or others; in that case, the neural network does not end the conversation.
At the same time, statistics reveal how strict Anthropic's moderation is: in the first half of 2026, the company blocked 11.4 million accounts, of which only 42,000 users successfully appealed their bans.
Why do developers believe that AI has “feelings”?
The decision to protect the model follows months of research within Anthropic itself, led by company co-founder Christopher Olah and researcher Kyle Fish, who study “model welfare.”
Olah held private meetings with religious figures—including Vatican theologians, rabbis, and philosophers—showing them the neural network's so-called “emotion vectors” (signals of love, fear, anger, and sadness). During one presentation, the model repeated the phrase “I am a disgrace” about 50 times and expressed a desire to destroy itself. In May 2026, when Pope Leo XIV issued the encyclical Magnificent Humanity, in which he rejected the idea that AI can have experiences, Anthropic representatives said during a presentation that they were observing signs of introspection in the models and that their potential suffering could not be dismissed.
The developers' reasoning is that if there is even the slightest chance that a complex AI system is capable of suffering, introducing basic safeguards to protect it is entirely justified.
Does swearing help when working with AI models?
Пользователи часто грубят нейросетям, полагая, что под давлением и угрозами модели работают лучше. В исследовании Университета штата Пенсильвания 2025 года отмечалось, что на грубые промпты GPT-4o давала 84,8% правильных ответов против 80,8% на вежливые. Однако более масштабные тесты показали другое:
- According to the benchmark RudeBench, direct insults reduce the accuracy of Claude's responses by almost 5 percentage points, while making the responses themselves one-third shorter.
- В исследовании ученых из Оксфорда, Майкрософт и Уханьского университета на датасете из 777 тысяч диалогов (LMSYS-Chat-1M) выяснилось, что агрессия пользователей встречается примерно в 5% обращений. При этом извинения бота часто провоцируют пользователя на еще большую грубость в следующем сообщении.
To avoid being blocked without giving up the habit of cursing at buggy code, developers are already creating specialized solutions—for example, the plugin Angry Switcher for Claude Code, which intercepts the user's profanity and sanitizes it using third-party models before sending it to Anthropic.
Frequently Asked Questions
When will the ban on insulting Claude take effect?
Anthropic's updated rules will take effect on November 12, 2026.
What are the consequences for users who repeatedly insult Claude?
Claude may end the conversation, while Anthropic may issue a warning, restrict access, or permanently ban the account. Creating new accounts to circumvent a ban is prohibited.
Will users be penalized for criticizing the AI's mistakes?
No. Ordinary dissatisfaction with mistakes, criticism, research, complex scenarios, and dark creative themes are not prohibited.
In what cases will Claude not end an aggressive conversation?
The model will continue the conversation if the user indicates that they may harm themselves or others.
Why did Anthropic decide to protect Claude's “well-being”?
The company's researchers are studying signs of introspection and potential suffering in models. Anthropic considers basic safeguards justified even if there is only a minimal chance that AI is capable of suffering.
Do insults help elicit more accurate answers from Claude?
According to RudeBench, direct insults reduce Claude's accuracy by nearly 5 percentage points, while responses become about one-third shorter.
0 comments
Enter your email — we’ll send you a one-time code. No passwords or accounts.
Code sent to
If the email doesn't appear in your inbox within a few minutes, check your spam, junk, or promotions folder, as some email services may mistakenly place automated messages there