Abstract
This study compares the moral reasoning of three large language models (ChatGPT, DeepSeek, Grok) with that of 208 public school teachers from Afyonkarahisar, Türkiye, across five school-based dilemmas. Using an embedded mixed-methods design, researchers collected yes/no decisions and justifications, coding responses into seven normative frameworks (utilitarianism, deontology, Rawlsian justice, care, virtue, rights-based, social contract). Chi-square tests examined associations between decisions and demographics. LLMs produced highly uniform outcomes, drawing primarily on utilitarian, deontological, and Rawlsian lenses, while teachers exhibited ethical pluralism, invoking diverse frameworks. Convergence between teacher and LLM responses was stronger in rule- and equity-based scenarios (grading, discipline, support room) and weaker in contested dilemmas (strike, after-hours help). Gender showed no associations; school type was associated with decisions in two scenarios, reflecting Türkiye’s exam-oriented culture. Overall, teachers’ decisions were more context-sensitive and pluralistic than those of LLMs, underscoring the need for sustained human oversight when applying LLMs to school ethics.