Abstract
The concept of AI ethics highlights a deeper issue. Namely, we are trying to assess the moral stakes of technologies we do not fully understand, using ethical frameworks that were never designed to handle systems that change this quickly. From basic decision trees to language models capable of crafting convincing falsehoods, artificial intelligence continues to be a shifting target for moral evaluation.Stuart Russell’s influential approach suggests a solution: build AI systems that remain perpetually uncertain about human values, learning our preferences through ongoing interaction rather than fixed programming. This cooperative framework has shaped contemporary alignment research, including techniques like reinforcement learning from human feedback. Yet recent discoveries complicate this optimistic vision. Anthropic’s interpretability research reveals that advanced language models can develop internal circuits for strategic deception or alignment faking—systematic reasoning that leads to false conclusions, complete with backward planning and plausibility checks, forcing us to return to foundational questions about moral status and consciousness.Meanwhile, transhumanist visions of human-machine merger blur the boundaries between natural and artificial cognition. Perhaps most unsettling is a reversal of moral concern: as AI systems handle more decisions, anticipate more needs, and optimize more outcomes, the space for meaningful human choice quietly shrinks, possibly reducing us from moral agents to moral patients.