Anthropic bans ‘abusive or cruel behavior’ towards Claude
The Verge·1 min read
Last August, the company announced it would allow Claude to end conversations with “persistently harmful or abusive” users as part of its research into “model welfare.” The new update says that terminating conversations is still the “primary enforcement mechanism”; Anthropic did not provide a comment on whether there would be further enforcement mechanisms, such as potential user bans. Anthropic wrote that the new policy update “is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.”