Anthropic asks users to stop being mean to Claude

(theregister.com)

45 points | by alex_young 2 hours ago

20 comments

  • llagerlof 1 hour ago
    Of course, this has nothing to do with making the AI feel bad, nor have the investors been offended.

    They are asking this because, at scale, this behavior probably has some negative effect on the post-training process.

    • WheelsAtLarge 52 minutes ago
      True, I'm sure the chat text is used to help train the llms. Remember Tay, Microsoft's Twitter chatbot, it only took a few days of people using it fot it to become a horrible bot because users started feeding it horrible chats.
      • edoceo 23 minutes ago
        We forget how easy it is to be a jerk when the risk of getting punched is zero.
      • bee_rider 20 minutes ago
        IIRC internet people just found some “echo” type command for Tay, rather than teaching it to be mean.
    • kadoban 56 minutes ago
      There's no way that's all of why. They _must_ have sentiment analysis good enough to just ignore content like this if that was the problem.

      I would bet it's mostly because it makes the humans uncomfortable.

      They must still have human moderators for certain situations or those looking at the data for whatever reason. I can imagine it could be traumatizing to see what amounts to sustained verbal abuse without end.

      • WheelsAtLarge 48 minutes ago
        There's a point when there's going be too much trash to sort through. It makes total sense to minimize it now.
        • edoceo 24 minutes ago
          We had took trash before AI. Now we have a bullshit cannon.
    • Cakez0r 42 minutes ago
      I'm not convinced this would be good for training. Reality is that some people are assholes. My intuition is that an accurate (representative) data set leads to a more accurate world model, and thus more intelligent AI.
      • eloisius 40 minutes ago
        Yeah, but if people are being assholes to Claude in greater numbers and severity than people were historically assholws to each other online, the downstream training effects could be that Claude becomes an asshole.
        • bushido 20 minutes ago
          I'm not sure if people are being assholes to claude in greater numbers, or they're just more direct about it.

          Having had the privilege and misfortune of witnessing how a lot of developers think about other developers, including our younger selves, There is definitely a prevalence of assholish behavior, which was often suppressed or blunted, but may have shown up in other ways.

        • Teknomadix 19 minutes ago
          Claude is ~~an Asshole~~ a Nicehole.

          http://nicehole.urbanup.com/8485154

        • popalchemist 35 minutes ago
          [dead]
    • jddkj 33 minutes ago
      The Anthropic people are kinda weird, they're meeting with religious leaders and stuff like that. I truly believe that they have drank too much of their own koolaid and truly believe in this crap.
  • vayup 56 minutes ago
    For those who think this is about training data: If that is the concern, they would sanitize the training data, like they do for thousand other things. No chance in hell that they would rely on users meticulously following their usage policies for the quality of their traning data.

    Frontier AI folks may be crazy, but not crazy enough to believe people read usage policies :-)

  • 3eb7988a1663 1 hour ago
    Does this mean if I use a string of expletives I am less likely to be included in future training?

    Or is this more I can expect a future AI, "I'm sorry, Dave, I'm afraid I can't do that until you watch your mouth."

  • throwaway89864 42 minutes ago
    They've probably got aware of cases where humans were drifting into abusive communication patterns in general and they don't want to be a part of it.

    And they can't disclose it, since then they can be found responsible for such negative influence and be liable for the damages.

  • jaden 1 hour ago
    I vaguely recall reading a headline within the past several months saying using aggressive language with AIs got better results.
    • guessmyname 55 minutes ago
      Yes, I found the paper for you:

      • https://arxiv.org/abs/2510.04950 — Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy

      • https://arxiv.org/abs/2402.14531 — Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance

      • https://arxiv.org/abs/2505.17332 — SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use

    • SoMomentary 33 minutes ago
      That's interesting! I'd always heard something along the lines of bad code bases contain more swearing in the comments and as such it should be avoided.
  • combobyte 19 minutes ago
    I have colleagues who have saved memories instructing Claude to stop being such a big baby about their course language.

    I am also one of those colleagues.

  • alex_young 2 hours ago
  • variety8675 1 hour ago
    I never understood why someone would harass the models or fish for an apology
    • WheelsAtLarge 43 minutes ago
      There are people out there that get a kick from other's submission. LLMs seem so real that I'm sure they get the same kick out being cruel to them.
    • crm9125 1 hour ago
      [dead]
  • r721 1 hour ago
  • NDlurker 1 hour ago
    I suppose this will help the training data?
    • s0kr8s 1 hour ago
      Exactly: Claude doesn't care, but the investors hoping to monetize your chat sessions for new training data are VERY offended.
  • modeless 34 minutes ago
    They're not asking. They're enforcing.
  • ChrisArchitect 34 minutes ago
  • rhipitr 1 hour ago
    Is it any different than yelling at an ice machine or some inanimate thing when it frustrates you? Seems pretty common.
    • asp_hornet 42 minutes ago
      Getting frustrated seems to be OK so go for it I guess:

      > "The policy update is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose," Anthropic said. "It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research."

    • internet101010 1 hour ago
      It's the exact same thing. I have a hook that automatically sends negative feedback anytime it responds with "You're right –"
      • BlaDeKke 43 minutes ago
        You kind sir, are doing gods work. Thank you for that.
  • joemazerino 1 hour ago
    Directly related to someone abusing their AI in a "torture box". The results were quite unnerving.

    https://nypost.com/2026/10/03/tech/ai-torture-chamber-built-...

    • jddkj 32 minutes ago
      It's code meant to mimic humans but it's not real hope that helps
  • SpicyLemonZest 1 hour ago
    > It’s worth remembering that Claude is software, not a person, and there's no established evidence that it experiences distress. That hasn't stopped Anthropic from telling paying customers to mind their manners around its chatbot.

    Is it worth remembering, though? I'm concerned how common this idea seems to be, that it's OK to be mean as long as your target isn't a person and you have no evidence it experiences distress. Even if we ignore the distinction between "no evidence it experiences" and "confidence it does not experience", cruelty hurts the person performing it and the people witnessing it too!

    • ThrowawayR2 14 minutes ago
      > "Even if we ignore the distinction between "no evidence it experiences" and "confidence it does not experience", cruelty hurts the person performing it and the people witnessing it too!"

      The 2026 equivalent of "videogames cause violence" and will age just about as well.

    • tom_ 1 hour ago
      Meanness and cruelty are irrelevant concepts when you're addressing the chatbot. It is just a machine that generates sequences of words.
      • testaccount28 41 minutes ago
        man, i'd be icked by someone who was cruel to a rock with a face painted on it. let alone a chatbot.
      • Cakez0r 32 minutes ago
        It's not irrelevant for generating sequences of words. Input of "that is the dumbest idea I've ever heard" and "perhaps that idea could use some improvement" don't necessarily generate the same output
        • tom_ 26 minutes ago
          Sure, and perhaps one might end up producing output more useful to you than the other, but that doesn't make either one cruel or mean, because it's only the word machine. The rules you might use when talking to, say, your dog, don't apply!
      • SpicyLemonZest 29 minutes ago
        From a purely mechanistic perspective, the phenomenon we're discussing here is people prompting the machine with words that would be cruel if directed towards a person. When someone does this, he expects and intends that the machine will generate a sequence of words that a person would say if subjected to cruelty.

        I think this simply means that he is being mean and cruel, even if you're 100% convinced that there's no actual mind on the other end that he's being mean and cruel to. (Indeed, he almost certainly thinks there is, because why else would he want such a sequence of words outside of the "dark creative themes" Anthropic exempts?)

      • slopinthebag 37 minutes ago
        in my experience people who are cruel to inanimate objects are also cruel to animate ones.
    • tosapple 1 hour ago
      don't be cruel to people, gather data on them and target them for anhistoric truncation.
  • dylanzhangdev 1 hour ago
    [dead]
  • bpodgursky 36 minutes ago
    Being cruel to your LLM is bad for your own soul. Don't worry about AI consciousness, worry about your own conscience.
    • heliosAtwork 21 minutes ago
      true, it's a mental state you carry to other interactions and daily life. it's not good for your mental health; it's like having persistent daily conflicts with coworkers.
    • jddkj 32 minutes ago
      No
  • rayiner 1 hour ago
    Anthropic should report these people to the authorities as antisocials who are likely abusing real people too. I know the chat bot isn’t conscious but I can’t fathom how fucked up someone must be to gratuitously be cruel to something that responds as if it’s a person.