The WuXing Architecture: Systems-Theoretic Foundations and Philosophical Principles of AI Self-Adversarial Systems

Abstract

The safety governance of large language models faces a fundamental impasse: patchwork defenses perpetually lag behind attackers' innovations, while end-to-end alignment approaches such as RLHF suffer from high costs, poor explainability, and difficulty covering long-tail scenarios. This paper proposes an AI safety architecture based on classical Chinese WuXing (Five Phase) theory — Wood, Fire, Earth, Metal, Water — which we term the "WuXing Architecture." The central insight is that safety is not a problem to be "solved," but a dynamic process that must be "navigated." The architecture decomposes an AI system into five functional layers connected through two fundamental relationships: "Generative Interaction" (mutual promotion) and "Restrictive Interaction" (mutual constraint), forming a closed-loop system capable of self-adversarial, self-repairing, and self-evolving behavior. Drawing on general systems theory, cybernetics, and complexity science, the paper demonstrates that genuine system stability arises not from absolute control but from dynamic equilibrium sustained by internal contradiction and counterbalancing forces. The WuXing Architecture offers a paradigm shift from fortress-like defense to ecosystem-like resilience, and introduces a non-Western systems framework into the foundations of AI safety philosophy.

Author's Profile

Analytics

Added to PP
2026-06-26

Downloads
608 (#88,329)

6 months
608 (#4,215)

Historical graph of downloads since first upload
This graph includes both downloads from PhilArchive and clicks on external links on PhilPapers.
How can I increase my downloads?