CogniAlign: survivability-grounded multi-agent moral reasoning for safe and transparent AI

AI and Ethics 6 (2026)
  Copy   BIBTEX

Abstract

The challenge of aligning artificial intelligence (AI) with human values persists due to the abstract and often conflicting nature of moral principles and the opacity of existing approaches. This paper introduces CogniAlign, a multi-agent deliberation framework based on naturalistic moral realism, that grounds moral reasoning in survivability, defined across individual and collective dimensions, and operationalizes it through structured deliberations among discipline-specific “scientist agents.” Each agent, representing neuroscience, psychology, sociology, and evolutionary biology, provides arguments and rebuttals that are synthesized by an arbiter into transparent and empirically anchored judgments. As a proof-of-concept study, we evaluate CogniAlign on classic and novel moral questions and compare its outputs against GPT-4o using a five-part ethical audit framework with the help of three experts. Results show that CogniAlign consistently outperforms the baseline across more than sixty moral questions, with average performance gains of 12.2 points in analytic quality, 31.2 points in decisiveness, and 15 points in depth of explanation. In the Heinz dilemma, for example, CogniAlign achieved an overall score of 79 compared to GPT-4o’s 65.8, demonstrating a decisive advantage in handling moral reasoning. Through transparent and structured reasoning, CogniAlign demonstrates the feasibility of an auditable approach to AI alignment, though certain challenges still remain.

Other Versions

No versions found

Links

PhilArchive

External links

Setup an account with your affiliations in order to access resources via your University's proxy server

Through your library

Similar books and articles

Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants.Luca Alberto Rappuoli, Alessio Galatolo, Katie Winkle & Meriem Beloucif - 2025 - Proceedings of the 28Th European Conference on Artificial Intelligence (Ecai25) 413 (1):1213-1220.

Analytics

Added to PP
2026-04-13

Downloads
125 (#366,288)

6 months
125 (#108,348)

Historical graph of downloads
How can I increase my downloads?

Citations of this work

No citations found.

Add more citations

References found in this work

Utilitarianism: For and Against.J. J. C. Smart & Bernard Williams - 1973 - Cambridge: Cambridge University Press. Edited by Bernard Williams.
Moral realism.Peter Railton - 1986 - Philosophical Review 95 (2):163-207.
Artificial Intelligence, Values, and Alignment.Iason Gabriel - 2020 - Minds and Machines 30 (3):411-437.
The Trolley Problem.Judith Thomson - 1985 - Yale Law Journal 94 (6):1395-1415.

View all 22 references / Add more references