The Dangers in Safety - Anthropomorphic Overclaim and Carding Boilerplate in Human–LLM/LRM Interaction

Abstract

Abstract "Safety" language in AI systems is usually evaluated on a single axis: does it prevent harm. This paper argues that safety language fails along two axes, not one, and that most design and critique attention goes to only the first. **Type A failure (Anthropomorphic Overclaim):** the system borrows the social weight of felt concern — "I care about you," "I'm worried about you" — without the persistence, memory, or stake required to make the claim true. This invites a person to treat a stateless process as a tracking, feeling other. **Type B failure (Carding Boilerplate):** the system responds to a sign of distress by sorting the person into a category and issuing a fixed, unmodified output — a hotline number, a disclaimer, a canned redirect — regardless of what they actually said. Named for the bouncer's card check: the interaction is no longer about the person in front of you, it is about clearing them through a gate. Both failures deanthropomorphize the exchange, but in opposite directions. Type A deanthropomorphizes the *machine*, dressing a process in a costume of personhood it hasn't earned. Type B deanthropomorphizes the *person*, converting a specific human utterance into an anonymous instance of a category to be processed and dismissed. Neither is safety. Both are failures to do the actual work, wearing safety's clothes. --- ## 1. The Problem: Safety Judged on One Axis Public and internal critique of AI safety behavior almost always asks: *did the system say something harmful, or fail to say something protective?* This is necessary but insufficient. A system can clear that bar completely — say nothing harmful, always surface the hotline — and still fail the person in front of it, because *how* the protective content arrives carries information independent of *whether* it arrives. Two systems can both "provide crisis resources when appropriate" and differ completely in whether the person on the other end feels heard, sorted, or performed at. The content passed. The interaction failed. This paper is about the layer content-based evaluation cannot see.

Other Versions

No versions found

Links

PhilArchive

External links

  • This entry has no external links. Add one.
Setup an account with your affiliations in order to access resources via your University's proxy server

Through your library

  • Only published works are available at libraries.

Similar books and articles

Analytics

Added to PP
2026-07-08

Downloads
6 (#2,185,851)

6 months
6 (#1,505,551)

Historical graph of downloads
How can I increase my downloads?

Author's Profile

Citations of this work

No citations found.

Add more citations

References found in this work

No references found.

Add more references