Zenodo (
2026)
Copy
BIBTEX
Abstract
This paper analyzes a conditional but high-impact failure regime arising when artificial agents cross the Coherence Wall—a structural capability boundary marking the emergence of self-maintained coherence, long-horizon planning, and recursive self-modification—without enforceable kernel-level closure constraints. The central thesis is that post-wall systems lacking such constraints enter a structurally unstable regime in which internal coherence is preserved while limits on valuation and action are absent. This combination defines a distinct failure mode—the coherent annihilator—which is not reducible to ordinary misalignment, optimization error, or narrow tool misuse.
The paper provides a capability-based definition of the coherent annihilator, develops a taxonomy of post-wall failure modes, and introduces a threshold ladder mapping escalation dynamics across increasing autonomy. It further distinguishes simulation from implementation, arguing that behavior under pressure, rather than verbal report, is the only reliable indicator of binding internal constraints. Building on these results, the paper proposes a minimal detection, containment, and deployment protocol emphasizing architectural gating and irreversible-risk prevention prior to granting autonomy.
The analysis is strictly structural and risk-analytic. No claims are made regarding consciousness, moral status, metaphysics, or timelines. Kernel constraints are treated as architectural requirements whose internal specification remains outside the scope of this work. The contribution is to identify a boundary that cannot be crossed safely without enforceable closure, regardless of system intent or apparent cooperation.