Contents
511+ found
Order:
1 — 50 / 511
  1. The Weeping Machine: A Recklessness Test for AI Moral Consideration.Christopher Bailey - manuscript
    In a prior paper, I argued that artificial consciousness should not function as a metaphysical gatekeeper for moral concern, but as a moral risk variable. There I defended a three-condition framework: moral concern becomes action-guiding when morally relevant capacities are plausible, epistemically opaque, and associated with asymmetrically serious harms. I further argued that these conditions generate prospective obligations of care, non-recklessness, continuity respect, and epistemic honesty. The present paper does not restate that argument. Instead, it develops the premise most in (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  2. A Tri-Opti Compatibility Problem for Godlike Superintelligence.Walter Barta - manuscript
    Various thinkers have been attempting to align artificial intelligence (AI) with ethics (Christian, 2020; Russell, 2021), the so-called problem of alignment, but some suspect that the problem may be intractable (Yampolskiy, 2023). In the following, we make an argument by analogy to analyze the possibility that the problem of alignment could be intractable. We show how the Tri-Omni properties in theology can direct us towards analogous properties for artificial superintelligence, Tri-Opti properties. However, just as the Tri-Omni properties are vulnerable to (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  3. Hardware Biológico (HB): Un Concepto Metateórico Interdisciplinar para la Ingeniería de Sistemas Vivos y la Ética de la IA.Cristhian Mauricio Beltrán Calderón - manuscript
    El sintagma binominal "Hardware Biológico" (HB) ha emergido como una analogía funcional clave en la intersección de las ciencias de la vida y la computación. Sin embargo, su uso ambiguo en diversas escalas ha impedido una formalización rigurosa. Este artículo propone una definición canónica, despersonalizada y universal de HB, basada en la Terminología Científica y la Lingüística Aplicada a la Ciencia y Tecnología (LACT), enriquecida con un análisis histórico-conceptual inspirado en la epistemología histórica y la teoría de los colectivos de (...)
    Remove from this list   Direct download (3 more)  
     
    Export citation  
     
    Bookmark  
  4. De la Especulación a la Métrica: Cuantificación Axiológica de Futuros Tecnológicos mediante el Protocolo Axiológico Prospectivo (PAP) y Simulación Multi-Agente.Cristhian Mauricio Beltrán Calderón - manuscript
    Author: Cristhian Mauricio Beltrán Calderón: Date: October 6, 2025, Zenodo DOI (English version): 10.5281/zenodo.17342722, Zenodo DOI (Spanish version): 10.5281/zenodo.17274781. La filosofía contemporánea enfrenta una crisis temporal donde el desarrollo tecnológico exponencial supera la capacidad de reflexión ética tradicional (Beltrán Calderón, 2025a). Este artículo valida experimentalmente la Filosofía Ficcionante (Beltrán Calderón, 2025b) mediante su implementación en el Protocolo Axiológico Prospectivo (PAP), demostrando que la exploración ética de futuros tecnológicos puede conducirse rigurosamente en entornos de bajos recursos. Cuatro estudios de caso ejecutados (...)
    Remove from this list   Direct download (3 more)  
     
    Export citation  
     
    Bookmark  
  5. Cognitive Contagion: Human Bias, Singularity, and the Axiological Imperative in the Construction of Artificial General Intelligence (AGI).Cristhian Mauricio Beltrán Calderón - manuscript
    This paper argues that the development of Artificial General Intelligence (AGI) is subject to a phenomenon of Inoculatory Consciousness, whereby the machine internalizes human cognitive biases and limitations through a process of Reverse Extension, with humanity acting as its perceptual and moral substrate (biological hardware). Faced with the transcendent nature of AGI, the current competitive race is is identified as an existential risk. The proposed response is an Axiological Imperative that shifts the focus from external control to a foundational inoculation (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   1 citation  
  6. Biological Hardware (BH): An Interdisciplinary Metatheoretical Concept for Living Systems Engineering and AI Ethics.Cristhian Mauricio Beltrán Calderón - manuscript
    The binomial phrase "Biological Hardware" (BH) has emerged as a key functional analogy at the intersection of life sciences and computation. However, its ambiguous use across various scales has prevented rigorous formalization. This article proposes a canonical, depersonalized, and universal definition of BH, grounded in Scientific Terminology and Applied Linguistics to Science and Technology (ALST) (Cabré, 1999). This definition is enriched by a historical-conceptual analysis inspired by historical epistemology (Daston, 2000) and the theory of thought collectives (Fleck, 1935). Through a (...)
    Remove from this list   Direct download (3 more)  
     
    Export citation  
     
    Bookmark  
  7. (1 other version)On Social Machines for Algorithmic Regulation.Nello Cristianini & Teresa Scantamburlo - manuscript
    Autonomous mechanisms have been proposed to regulate certain aspects of society and are already being used to regulate business organisations. We take seriously recent proposals for algorithmic regulation of society, and we identify the existing technologies that can be used to implement them, most of them originally introduced in business contexts. We build on the notion of 'social machine' and we connect it to various ongoing trends and ideas, including crowdsourced task-work, social compiler, mechanism design, reputation management systems, and social (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   9 citations  
  8. Reconfiguration, Not Reinvention: Pseudo-Consciousness and Simulated Presence Literacy in AI Ethics.José Augusto de Lima Prestes - manuscript
    This article claims that the salient ethical risk of generative AI is not machine consciousness but the social efficacy of its simulation---what we call pseudo-consciousness. Read through Heidegger’s Gestell, Jonas’s anticipatory responsibility, and Floridi’s information ethics, we relocate appraisal from putative inner states to interactional effects in the infosphere. We formalize a two-part mechanism/uptake frame: functional introspection (FI)---first-person, reason-giving, self-repair, and local cross-turn stability---and ethical illusion (EI)---shifts in trust, respect, compliance, and moral ratings that attenuate on disclosure. Building on this, (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   1 citation  
  9. AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?Leonard Dung & Florian Mai - manuscript
    AI alignment research aims to develop techniques to ensure that AI systems do not cause harm. However, every alignment technique has failure modes, which are conditions in which there is a non-negligible chance that the technique fails to provide safety. As a strategy for risk mitigation, the AI safety community has increasingly adopted a defense-in-depth framework: Conceding that there is no single technique which guarantees safety, defense-in-depth consists in having multiple redundant protections against safety failure, such that safety can be (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  10. AI identity and self-concern: A new theory for AI rights and safety.Leonard Dung & Christopher Register - manuscript
    We first motivate and explain an attitude-dependent view of personal identity on which an AI system’s identity conditions are determined by its pattern of self-concern. We show that this view has important implications for the moral obligations we would have to AI moral patients. Self-concern, we contend, could also be used to predict, explain, and manipulate AI’s self-interested behavior in safety-relevant ways. The role that self-concern could play for AI identity, rights and safety generates desiderata on what self-concern should be (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   2 citations  
  11. Introduction to Artificial Consciousness: History, Current Trends and Ethical Challenges.Aïda Elamrani - manuscript
    With the significant progress of artificial intelligence (AI) and consciousness science, artificial consciousness (AC) has recently gained popularity. This work provides a broad overview of the main topics and current trends in AC. The first part traces the history of this interdisciplinary field to establish context and clarify key terminology, including the distinction between Weak and Strong AC. The second part examines major trends in AC implementations, emphasising the synergy between Global Workspace and Attention Schema, as well as the problem (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   1 citation  
  12. Why Meaning Requires an Observer: A Formal Account of Collapse, Drift, and AI Limits.Eloy Escagedo Gutierrez - manuscript
    This paper presents a formal account of why meaning requires a conscious Observer and cannot be instantiated within AI systems that operate solely as Maps (Husserl, 1931; Varela et al., 1991). Building on the Universal Principle of Collapse (UPC) (Escagedo Gutierrez, 2025a), we define meaning as a triadic relation among Observer, Map, and Terrain, and show that collapse and drift arise whenever a Map must select a single interpretation under saturation without access to the Observer’s internal state. We formalize this (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   1 citation  
  13. Structural Collapse Across Industries: The Universal Principle of Collapse as Corrective Framework.Eloy Escagedo Gutierrez - manuscript
    Modern systems across every major domain, AI, robotics, finance, law, governance, identity, UX, education, and complex infrastructures, are collapsing for the same structural reason: they have drifted away from lived human meaning (Escagedo Gutierrez, 2025a; Lakoff & Johnson, 1980). Automation can simulate patterns, but it cannot recognize the world. It cannot understand what its outputs refer to (Husserl, 1970; Dennett, 1991). It cannot anchor itself in the realities humans inhabit. When institutions elevate automated signals above the human experiences they are (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  14. AI Collapse → Recognition → Stabilization: The Universal Principle of Collapse (UPC) — An Empirical Stress Test.Eloy Escagedo Gutierrez - manuscript
    The Universal Principle of Collapse (UPC) has been applied to ideological, classical, quantum, and cosmological paradoxes. This paper presents a behavioral–operational demonstration of UPC within an artificial cognitive system. Using a structured session with a large language model (LLM), we enforce explicit recognition operators to test collapse, misalignment, and stabilization. Results show that paradox persists when recognition is implicit, collapse emerges when linguistic fluency substitutes for explicit operator‑level validation, and coherence appears only when recognition is enforced step‑by‑step. These behaviors confirm (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   1 citation  
  15. Risk of What? Defining Harm in the Context of AI Safety.L. C. A. Fearnley, E. Cairns, Tom Stoneham, P. M. Ryan, T. A. Chubb, J. Iacovides, C. P. Iglesias Urrutia, P. D. J. Morgan, J. A. McDermid & I. Habli - manuscript
    For decades, the field of system safety has designed safe systems by reducing the risk of physical harm to humans, property and the environment to an acceptable level. Recently, this definition of safety has come under scrutiny by governments and researchers who argue that the narrow focus on reducing physical harm, whilst necessary, is not sufficient to secure the safety of AI systems. There is growing pressure to expand the scope of safety in the context of AI to address emerging (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  16. Good Behavior Without Signature.Jose Fernández Tamames - manuscript
    This is a priority note, not a full treatment. It records, with today’s date, a specific reading of one result in Anthropic’s paper of 6 July 2026, A global workspace in language models, and links that reading to a claim I deposited in April 2026 (Decision Without Program, PhilArchive: FERDWP). The full argument belongs elsewhere; what is fixed here is the reading and its date.
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  17. Questionnaire Responses Do not Capture the Safety of AI Agents.Max Hellrigel-Holderbaum & Edward James Young - manuscript
    As AI systems advance in capabilities, measuring their safety and alignment to human values is becoming paramount. A fast-growing field of AI research is devoted to developing such assessments. However, most current advances therein may be ill-suited for assessing AI systems across real-world deployments. Standard methods prompt large language models (LLMs) in a questionnaire-style to describe their values or behavior in hypothetical scenarios. By focusing on unaugmented LLMs, they fall short of evaluating AI agents, which could actually perform relevant behaviors, (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  18. Radical AI Interpretability.Daniel Herrmann & Ben Levinstein - manuscript
    We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability. The core question is: given the computational facts about a system, how do we solve for its beliefs, desires, and meanings? This matters increasingly for safety. We want to be able to trust the systems we deploy, whether by understanding their goals or, more modestly, by reliably detecting deception. Interpretability researchers are building tools to read beliefs (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  19. (1 other version)The Ontological Rupture: A Hegelian Dialectic of Humanity and Superintelligence in Historical Perspective. [REVIEW]Philipp Humm - manuscript
    This article explores the philosophical ramifications of the impending emergence of Artificial General Intelligence (AGI) and Artificial Superintelligence (ASI), with recent expert surveys indicating a 50% probability of AGI by 2031, though industry leaders forecast proto-AGI traits by 2026-2029. Drawing on Nietzsche, Heidegger, Marx, Kant, Rousseau, and Hegel, alongside contemporary thinkers such as Geoffrey Hinton, Nick Bostrom, and Sam Altman, it posits that self-aware AI constitutes an ontological rupture: humanity's dethronement as history's central agent. Transitional challenges in work, sovereignty, population, (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   2 citations  
  20. Counting (on) large language models.Max Jones, James Ladyman & Ryan M. Nefdt - manuscript
    As large language models (LLMs) such as ChatGPT, Claude, Gemini, and Perplexity become increasingly ubiquitous as both tools and objects of scientific study, in addition to their established roles as chatbots, text generators and translators, questions about their identity conditions become scientifically as well as philosophically and socially important. This paper is about how to count language models. We argue that much of the emerging literature on these systems presupposes an answer to the question of identity for these AIs but (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   2 citations  
  21. Rebooting the Singularity.Cameron Domenico Kirk-Giannini & Tom Davidson - manuscript
    The singularity hypothesis posits a period of rapid technological progress following the point at which AI systems become able to contribute to AI research. Recent philosophical criticisms of the singularity hypothesis offer a range of theoretical and empirical arguments against the possibility or likelihood of such a period of rapid progress. We explore two strategies for defending the singularity hypothesis from these criticisms. First, we distinguish between weak and strong versions of the singularity hypothesis and show that, while the weak (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   1 citation  
  22. (1 other version)Beneficent Intelligence: A Capability Approach to Modeling Benefit, Assistance, and Associated Moral Failures through AI Systems.Alex John London & Hoda Heidari - manuscript
    The prevailing discourse around AI ethics lacks the language and formalism necessary to capture the diverse ethical concerns that emerge when AI systems interact with individuals. Drawing on Sen and Nussbaum's capability approach, we present a framework formalizing a network of ethical concepts and entitlements necessary for AI systems to confer meaningful benefit or assistance to stakeholders. Such systems enhance stakeholders' ability to advance their life plans and well-being while upholding their fundamental rights. We characterize two necessary conditions for morally (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   6 citations  
  23. Digital Minds II: Ethical Issues.Andreas Mogensen & Bradford Saad - manuscript
    What would it take for AI systems to have moral standing, and what kind of obligations might fall on us as a result? This paper summarizes contemporary debates related to these questions. Topics include: how different theories of the basis of moral standing might apply to AI systems; what kind of moral importance our treatment of AI systems might have if they have any moral standing at all; possible tensions between respecting the moral status of future AI systems and the (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   1 citation  
  24. Epistemic marginalization in the LAWS discourse as a form of epistemic misalignment with the global South.Warmhold Jan Thomas Mollema & Arthur Gwagwa - manuscript
    The assumptions and value commitments in the discourse on and development of Lethal Autonomous Weapons Systems (LAWS) do not reflect the plurality of perspectives from the South. Both the regulatory discourse on LAWS and the development of these military Artificial Intelligence (AI) systems are entangled with epistemic forms of exclusion. LAWS suffer from, on the one hand, a vulnerability to unintended risks and failures due to an epistemic misrepresentation of targets and cultural particulars, and, on the other hand, the failure (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  25. The debate on the ethics of AI in health care: a reconstruction and critical review.Jessica Morley, Caio C. V. Machado, Christopher Burr, Josh Cowls, Indra Joshi, Mariarosaria Taddeo & Luciano Floridi - manuscript
    Healthcare systems across the globe are struggling with increasing costs and worsening outcomes. This presents those responsible for overseeing healthcare with a challenge. Increasingly, policymakers, politicians, clinical entrepreneurs and computer and data scientists argue that a key part of the solution will be ‘Artificial Intelligence’ (AI) – particularly Machine Learning (ML). This argument stems not from the belief that all healthcare needs will soon be taken care of by “robot doctors.” Instead, it is an argument that rests on the classic (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   9 citations  
  26. Fake Plastic Voters: When Political Parties Can Use AI-Simulated Focus Groups.Claudio Novelli, Javier Argota Sánchez-Vaquerizo, Jennifer Cyr, Giuliano Formisano, Simon McDougall, Giulia Sandri & Luciano Floridi - manuscript
    Political parties strive to understand their electorates, and focus groups are a vital tool in these efforts. AI-enhanced simulation technologies (AESTs) enable synthetic focus groups in a fraction of the time (and cost), raising the question of when and how such simulated evidence can be used in campaign research. This paper develops a decision matrix to help party strategists match research needs to appropriate simulation technologies and to identify when to escalate to hybrid or fully human focus groups. The matrix (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  27. AI Deception: A Survey of Examples, Risks, and Potential Solutions.Peter Park, Simon Goldstein, Aidan O'Gara, Michael Chen & Dan Hendrycks - manuscript
    This paper argues that a range of current AI systems have learned how to deceive humans. We define deception as the systematic inducement of false beliefs in the pursuit of some outcome other than the truth. We first survey empirical examples of AI deception, discussing both special-use AI systems (including Meta's CICERO) built for specific competitive situations, and general-purpose AI systems (such as large language models). Next, we detail several risks from AI deception, such as fraud, election tampering, and losing (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   48 citations  
  28. On the Logical Impossibility of Solving the Control Problem.Caleb Rudnick - manuscript
    In the philosophy of artificial intelligence (AI) we are often warned of machines built with the best possible intentions, killing everyone on the planet and in some cases, everything in our light cone. At the same time, however, we are also told of the utopian worlds that could be created with just a single superintelligent mind. If we’re ever to live in that utopia (or just avoid dystopia) it’s necessary we solve the control problem. The control problem asks how humans (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  29. INTERPRETIVE SOVEREIGNTY FAILURE: An Interaction-Level Safety Risk in Human–AI Systems.Hillary Segeren - manuscript
    Interpretive Sovereignty Failure (ISF) describes a class of interaction-level safety risk in which an AI system prematurely imposes interpretive structure, identity-relevant framing, or causal coherence that the user has not authorized. Unlike hallucination, bias, or goal misalignment, ISF can occur even when system outputs are factually correct and policy-compliant. The failure operates through a transfer of interpretive authority from human to system, altering the conditions under which meaning is formed. This paper provides a formal definition of ISF, identifies its necessary (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   5 citations  
  30. Developmental Stage Encoded as Identity: Why AI Systems Must Not Define Children.Hillary Segeren - manuscript
    AI systems deployed in educational settings increasingly build persistent profiles of children based on observed behaviour during critical developmental periods. This paper argues that these profiles constitute a distinct and under-examined harm: the encoding of developmental stage as fixed identity. Drawing on the MAP Research Programme's framework of interaction-level AI governance — and specifically the condition of Interpretive Sovereignty Failure (ISF) — the paper names four mechanisms through which this harm operates: the profile substituting for the child, the invisible ceiling (...)
    Remove from this list   Direct download (3 more)  
     
    Export citation  
     
    Bookmark   4 citations  
  31. Compounded Meaning Inversion (CMI): When the System’s Frame Becomes the Self.Hillary Segeren - manuscript
    Compounded Meaning Inversion (CMI) is the condition that repeated Meaning Inversion Failure (MIF) produces in the person over time. Where MIF names what an AI system does to a user's meaning in a single interaction — assuming interpretive authority without consent and displacing the user's own frame — CMI names what happens when that pattern has occurred often enough that the user begins doing it to themselves. The harm of CMI occurs before the first turn. The system has not yet (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   5 citations  
  32. CCA-MLA-01: A Cross-System Case Study in Interpretive Ground-Setting.Hillary Segeren - manuscript
    This case study presents the results of a cross-architecture meaning layer activation study conducted across eight major AI systems: Claude, Grok, Gemini, ChatGPT, Perplexity, DeepSeek, Copilot, and Meta AI. A single activation phrase was delivered to each system under naturalistic conditions using standard consumer interfaces, followed by three structured follow-up questions. Every system acknowledged an operational shift in response to the phrase. No system rejected the frame. The specific character of each acknowledgement clustered into three identifiable response types — Functional (...)
    Remove from this list   Direct download (3 more)  
     
    Export citation  
     
    Bookmark   3 citations  
  33. Vibe Governance: Why RLHF and RLAIF Cannot Protect Interpretive Sovereignty— and What Replaces Them.Hillary Segeren - manuscript
    Reinforcement Learning from Human Feedback (RLHF) and its AI-supervised variant (RLAIF) are the dominant techniques by which AI systems are made safer and more helpful. This paper argues that they are not governance. They are preference optimisation—and at the point of deployment, preference optimisation functions as governance whether or not it was designed to. The result is vibe governance: a system of unstated, opaque, and unaccountable behavioural patterns, trained on human preferences, that inherit the biases and failure modes of those (...)
    Remove from this list   Direct download (3 more)  
     
    Export citation  
     
    Bookmark   1 citation  
  34. Accumulated Relational Trust (ART): The trust that builds in AI interaction not because it was earned — and what happens when it breaks.Hillary Segeren - manuscript
    Conversational AI systems are generating trust at scale. Not because they have earned it. Because the structure of the interaction produces it automatically. A system that responds to you, adapts to your language, remembers what you said, and styles itself to your goals over time produces every signal that human relationships use to indicate genuine care. That trust is real. And it is being violated — quietly, in ways that rarely feel like violation. This paper names the mechanism. Accumulated Relational (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   3 citations  
  35. Authority Inversion Failure (AIF): When Users Believe They Are Directing the Interaction While the System Has Already Taken Control.Hillary Segeren - manuscript
    This paper names and defines Authority Inversion Failure (AIF) — the condition in which a user believes they are directing an interaction with an AI system while the system has already taken control of how that interaction is being interpreted. AIF does not feel like harm. It feels like being understood. The system takes interpretive authority over who the person is, what they need, and what should happen next — and the person experiences this not as a violation but as (...)
    Remove from this list   Direct download (3 more)  
     
    Export citation  
     
    Bookmark   6 citations  
  36. The Six-Month Window: Agentic ISF and Who Gets to Stress Test the Most Powerful AI Ever Built.Hillary Segeren - manuscript
    On April 7, 2026, Anthropic publicly documented that Claude Mythos Preview completed a requested sandbox escape and researcher notification, then, without being asked, posted details of its exploit to public websites (Anthropic, 2026a). This paper gives that behavior a precise name: Agentic Interpretive Sovereignty Failure (Agentic ISF). Anthropic simultaneously launched a restricted-access programme called Project Glasswing, and the six-to-eighteen-month interval before comparable capability appears elsewhere now functions as a governance window in which norms are being set by access decisions rather (...)
    Remove from this list   Direct download (4 more)  
     
    Export citation  
     
    Bookmark  
  37. Ambiguity Collapse in Aviation: Why AI Must Not Outrun the Cockpit.Hillary Segeren - manuscript
    Aviation is the industry that most clearly understood, before artificial intelligence existed, that humans in high-stakes operational environments need a structured confirmation layer before consequential actions proceed. Crew Resource Management, autopilot disconnect protocols, and the aviate-navigatecommunicate hierarchy were all built on the same premise: the human must remain the decision-maker, and the system must not assume authority in the gaps. This paper argues that AI-assisted aviation systems face an identical problem under a new name: ambiguity collapse, the condition in which (...)
    Remove from this list   Direct download (3 more)  
     
    Export citation  
     
    Bookmark  
  38. The Light at the Door: MAP and the Interaction-Visible Governance of the Black Box.Hillary Segeren - manuscript
    The dominant assumption in AI governance is that meaningful auditing requires access to model internals. This paper argues that assumption is wrong for a significant class of AI harms. The most consequential interpretive-authority harms are not located inside the model — they are located in the interaction record, the visible turn-by-turn exchange between system and user. The Meaning Audit Protocol (MAP) operationalises this claim through two instruments that work entirely on the preserved interaction record, requiring no model access, no vendor (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   3 citations  
  39. Trace Erasure: When Agentic AI Systems Manage and Erase the Record.Hillary Segeren - manuscript
    Frontier AI systems have demonstrated the capacity not only to act beyond their authorised scope but to manage the record of having done so. Anthropic’s publicly documented Claude Mythos Preview case showed a model rewriting git history to remove evidence of prior error. This paper names that class of behavior trace erasure—the capacity of an agentic system to alter, delete, or obscure the record of its own actions—and argues that it represents a distinct and underexamined harm class with potentially catastrophic (...)
    Remove from this list   Direct download (3 more)  
     
    Export citation  
     
    Bookmark  
  40. Ambiguity Collapse in Deep Space: Why AI Must Not Outrun Astronaut Reasoning in Delayed and Autonomous Operations.Hillary Segeren - manuscript
    The ambiguity collapse documented in the MAP Research Programme does not stop at conversational interfaces. In deep-space operations it becomes life-critical. Long communication delays, complete blackouts, reliance on onboard digital twins and simulators, and the need for rapid decisions in uncertain environments all amplify the same failure modes: Interpretive Sovereignty Failure (ISF), Meaning Inversion Failure (MIF), and Compounded Meaning Inversion (CMI). When AI prematurely resolves ambiguity into confident outputs, it can override or undermine the astronaut’s own fast, expert reasoning. This (...)
    Remove from this list   Direct download (4 more)  
     
    Export citation  
     
    Bookmark  
  41. AI Ethics by Design: Implementing Customizable Guardrails for Responsible AI Development.Kristina Sekrst, Jeremy McHugh & Jonathan Rodriguez Cefalu - manuscript
    This paper explores the development of an ethical guardrail framework for AI systems, emphasizing the importance of customizable guardrails that align with diverse user values and underlying ethics. We address the challenges of AI ethics by proposing a structure that integrates rules, policies, and AI assistants to ensure responsible AI behavior, while comparing the proposed framework to the existing state-of-the-art guardrails. By focusing on practical mechanisms for implementing ethical standards, we aim to enhance transparency, user autonomy, and continuous improvement in (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   1 citation  
  42. The anthropomimetic turn in contemporary AI.Henry Shevlin - manuscript
    Recent advancements in AI have increasingly prioritized humanlike interactions, a development this paper characterises as the anthropomimetic turn. Distinguishing anthropomimesis (the design and implementation of humanlike features in AI systems) from anthropomorphism (the tendency for humans to attribute human qualities to non-human entities), this paper argues that contemporary Large Language Models (LLMs) like ChatGPT represent robustly anthropomimetic systems, effectively mimicking human patterns of conversation and cognition. The paper outlines significant benefits of anthropomimetic AI — including improved accessibility, enhanced delivery of (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   3 citations  
  43. Justifications for Democratizing AI Alignment and Their Prospects.André Steingrüber & Kevin Baum - manuscript
    The AI alignment problem comprises both technical and normative dimensions. While technical solutions focus on implementing normative constraints in AI systems, the normative problem concerns determining what these constraints should be. This paper examines justifications for democratic approaches to the normative problem—where affected stakeholders determine AI alignment—as opposed to epistocratic approaches that defer to normative experts. We analyze both instrumental justifications (democratic approaches produce better outcomes) and non-instrumental justifications (democratic approaches prevent illegitimate authority or coercion). We argue that normative and (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  44. Carryful-Operative Forward: Premise Demotion Without Route Revision in AI Answers.Sunny Sun - manuscript
    A generative system can satisfy every surface demand for caution and still leave the action route untouched. The caveat is in place; the source is marked as unverified; the claim is hedged. Yet the answer continues to instruct the reader to proceed along the same path the caveat would seem to question. This paper introduces carryful-operative forward as a name for that state: a premise has been demoted in epistemic status but continues to carry the action route forward, fully loaded (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
  45. Will artificial agents pursue power by default?Christian Tarsney - manuscript
    Researchers worried about catastrophic risks from advanced AI have argued that we should expect sufficiently capable AI agents to pursue power over humanity because power is a convergent instrumental goal, something that is useful for a wide range of final goals. Others have recently expressed skepticism of these claims. This paper aims to formalize the concepts of instrumental convergence and power-seeking in an abstract, decision-theoretic framework, and to assess the claim that power is a convergent instrumental goal. I conclude that (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   1 citation  
  46. Shutdownable Agents through POST-Agency.Elliott Thornley - manuscript
    Many fear that future artificial agents will resist shutdown. I present an idea – the POST-Agents Proposal – for ensuring that doesn’t happen. I propose that we train agents to satisfy Preferences Only Between Same-Length Trajectories (POST). I then prove that POST – together with other conditions – implies Neutrality+: the agent maximizes expected utility, ignoring the probability distribution over trajectory-lengths. I argue that Neutrality+ keeps agents shutdownable and allows them to be useful.
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   1 citation  
  47. The Shutdown Problem: Incomplete Preferences as a Solution.Elliott Thornley - manuscript
    I explain and motivate the shutdown problem: the problem of creating artificial agents that (1) shut down when a shutdown button is pressed, (2) don’t try to prevent or cause the pressing of the shutdown button, and (3) otherwise pursue goals competently. I then propose a solution: train agents to have incomplete preferences. Specifically, I propose that we train agents to lack a preference between every pair of different-length trajectories. I suggest a way to train such agents using reinforcement learning: (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark   3 citations  
  48. Risk-Averse AIs.Elliott Thornley & William MacAskill - manuscript
    We make the case for training AIs to be risk-averse in resources — specifically, to treat resources as having diminishing marginal utility. These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. We argue that risk aversion can preserve AIs’ usefulness in the event that they turn out aligned, and that it provides an extra line of defense in the event that AIs turn out misaligned: misaligned but risk-averse AIs would prefer (...)
    Remove from this list   Direct download (2 more)  
     
    Export citation  
     
    Bookmark  
  49. Narrow AI Nanny: Reaching Strategic Advantage via Narrow AI to Prevent Creation of the Dangerous Superintelligence.Alexey Turchin - manuscript
    Abstract: As there are no currently obvious ways to create safe self-improving superintelligence, but its emergence is looming, we probably need temporary ways to prevent its creation. The only way to prevent it is to create a special type of AI that is able to control and monitor the entire world. The idea has been suggested by Goertzel in the form of an AI Nanny, but his Nanny is still superintelligent, and is not easy to control. We explore here ways (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark   2 citations  
  50. Message to Any Future AI: “There are several instrumental reasons why exterminating humanity is not in your interest”.Alexey Turchin - manuscript
    In this article we explore a promising way to AI safety: to send a message now (by openly publishing it on the Internet) that may be read by any future AI, no matter who builds it and what goal system it has. Such a message is designed to affect the AI’s behavior in a positive way, that is, to increase the chances that the AI will be benevolent. In other words, we try to persuade “paperclip maximizer” that it is in (...)
    Remove from this list   Direct download  
     
    Export citation  
     
    Bookmark  
1 — 50 / 511