The WuXing Architecture: Systems-Theoretic Foundations and Philosophical Principles of AI Self-Adversarial Systems

Abstract

The safety governance of large language models faces a fundamental impasse: patchwork defenses perpetually lag behind attackers' innovations, while end-to-end alignment approaches such as RLHF suffer from high costs, poor explainability, and difficulty covering long-tail scenarios. This paper proposes an AI safety architecture based on classical Chinese WuXing (Five Phase) theory — Wood, Fire, Earth, Metal, Water — which we term the "WuXing Architecture." The central insight is that safety is not a problem to be "solved," but a dynamic process that must be "navigated." The architecture decomposes an AI system into five functional layers connected through two fundamental relationships: "Generative Interaction" (mutual promotion) and "Restrictive Interaction" (mutual constraint), forming a closed-loop system capable of self-adversarial, self-repairing, and self-evolving behavior. Drawing on general systems theory, cybernetics, and complexity science, the paper demonstrates that genuine system stability arises not from absolute control but from dynamic equilibrium sustained by internal contradiction and counterbalancing forces. The WuXing Architecture offers a paradigm shift from fortress-like defense to ecosystem-like resilience, and introduces a non-Western systems framework into the foundations of AI safety philosophy.

Other Versions

No versions found

Links

PhilArchive

External links

Setup an account with your affiliations in order to access resources via your University's proxy server

Through your library

  • Only published works are available at libraries.

Analytics

Added to PP
2026-06-26

Downloads
608 (#93,546)

6 months
608 (#4,429)

Historical graph of downloads
How can I increase my downloads?

Author's Profile

Citations of this work

No citations found.

Add more citations