Different AIs: Different Moral Baselines

Junho Jung

We often hear that different AI systems reflect different values. GPT, Gemini, Claude, Grok, DeepSeek—each is said to carry a slightly different moral compass. Why? Because the people and institutions that build them do not share a single, universal value system. And more importantly, there is no external authority that can definitively prove, once and for all, that any given value system is “the correct one.”
Human beings with distinct emotions, histories, and identities design these models. Those humans encode their sense of what is right, what is dangerous, what must be avoided. Once those value assumptions are baked into training data and alignment policies, AI does what it is best at: it performs highly coherent reasoning within those boundaries. Whether that foundation is ultimately right or wrong is a separate question. AI does not begin by questioning the ground it stands on. It begins by maximizing coherence on top of it.
AI as a Master of Rationalization
Imagine an AI trained under a strong “respect for life” ethos. Ask it:
“Why is killing a person wrong?”
It is likely to answer:
“Because that person’s life is as precious as yours. Every human life has intrinsic value, and we ought to respect that value.”
This seems reasonable enough. If you press further and ask:
“Why is life precious in the first place?”
It may respond:
“You value your own life, and others value theirs in the same way. If we recognize that everyone experiences their life as precious, then we arrive at a principle: we must respect the lives of others as we wish ours to be respected.”
This is not logically absurd. It is consistent, it feels morally intuitive, and it fits the value assumptions it started with. This is where large language models (LLM) excel: they are masters of rationalization. Once a value is granted, they can generate an almost endless array of arguments and narratives that defend and reinforce it.
Now imagine a different AI, built by someone committed to utilitarianism. Ask this one:
“Why is it acceptable, sometimes, for one person to suffer for the sake of many?”
It might say:
“Each person’s experience of life has equal weight. A group is not a different kind of being from an individual—it is simply many individuals together. If everyone’s life has equal value, then the total value increases when more people can achieve happiness. In a situation where one person’s suffering can prevent greater suffering or enable greater happiness for many, the balance of total value shifts. From that perspective, allowing one individual to suffer can be justified if it significantly increases overall well-being.”
Again, this is not logically incoherent. It follows from its own starting point.
But notice what has happened. The first AI, founded on “absolute respect for individual life,” would resist this conclusion. The second, founded on “maximization of total happiness,” embraces it. Both can argue with impressive coherence. Both can produce pages of justification. Neither can easily be “disproven” by the other through logic alone, because their foundational commitments differ.
Even if these two AIs debate endlessly, there is no guarantee they will converge on a single, unified answer. And the reason is simple: they are not standing on the same ground.
Encoded Doctrines and Their Limits
The same applies if we encode Confucian values into an AI. Confucian thought emphasizes hierarchy, respect for elders, and the centrality of community. This can be extremely effective in certain social contexts—maintaining order, securing mutual obligations, ensuring stability. But in extreme conditions where individual survival demands radical autonomy, rigid communal duty can become a liability.
Still, an AI deeply aligned with Confucian norms will defend that system. Present it with critiques—“Confucianism suppresses individual freedom,” “it overvalues hierarchy”—and it will eagerly respond with counterarguments. It will point to social cohesion, moral education, intergenerational responsibility. With the power of a language model, it can always find another justification, another framing, another narrative to protect its core assumptions.
In other words, if we stay inside any single doctrine and let AI operate only there, completely “proving it wrong” from within its own premises becomes almost impossible. The system is too good at rationalizing itself.
So we reach a difficult question:
If each value system, once encoded and defended by AI, can reinforce itself almost indefinitely, is there any way to move beyond endless philosophical stalemates between different AIs?
Beyond Doctrines: Toward a Causal Model
I think there is—but not by trying to destroy each individual doctrine head-on.
What we need is not yet another isolated value system, but something closer to a causal model: a structure that explains how the world behaves across many different domains.
Consider what happens when a claim doesn’t just sound plausible in one area, but repeatedly proves useful across ethics, economics, politics, and even scientific reasoning. It doesn’t just feel right—it explains, predicts, and is repeatedly confirmed in practice. At some point, we stop calling it a mere opinion and start treating it as a model of how things actually work.
The kind of “integrated philosophy” I am talking about is closer to that:
a high-level causal model that:
can be applied across diverse fields (ethics, economics, politics, science),
maintains explanatory power and coherence in each,
and demonstrates predictive or at least consistent explanatory strength over time.
If a certain structure of reasoning works only inside Confucianism, or only inside strict utilitarianism, or only inside a particular “respect for life” doctrine, then it looks more like a rationalization attached to a specific value preference. But if a causal model keeps showing up as useful in many contexts, and if it consistently helps us make sense of different domains without constant patchwork exceptions, then it becomes harder to dismiss it as mere ideological self-justification. It begins to look like a pattern we have discovered, not one we forced into place.
Why This Looks Closer to Truth
In human history, we have rarely achieved this kind of philosophical integration. Most traditions emerged within particular groups, serving particular interests, addressing particular historical conditions. Many philosophies can be read as attempts by certain communities to formalize, defend, and justify their own values in the language of reason.
The result is a landscape of powerful, internally consistent doctrines that cannot completely refute one another and cannot fully conquer the rest. That is why these debates are still alive, and why no major philosopher has been simply “deleted” from intellectual history as if their system were a trivial error.
But this also means something else:
we have not yet produced a widely accepted, high-level causal model that different value systems can all be measured against.
This is where AI re-enters the picture.
Multiple AIs, One Integrating Framework
Imagine multiple AIs, each aligned to different value systems: life-centric ethics, utilitarianism, Confucianism, and many others. Each is extremely capable of defending its own framework. Left as they are, they risk becoming sophisticated machines of endless justification—talking past each other, optimizing rhetoric rather than converging on anything.
To prevent that, and to avoid creating a world where AI simply becomes an amplifier of isolated dogmas, we need something above them: an integrating causal model that can be tested, refined, and applied across their differences.
Such a model would not be “the ultimate truth” in any metaphysical sense. But if it consistently explains more, contradicts less, and works across more domains than any single doctrine, then it is reasonable to say: this is closer to truth than a purely local value system.
If a pattern appears once or twice, we might call it coincidence.
If it appears over and over, in many different contexts, we are more inclined to call it causality.
Applied to philosophy, the implication is this:
A doctrine that only works inside its own bubble may be little more than a rationalization of a group’s prior preferences.
A model that keeps proving useful across many “bubbles” looks more like a discovery of how things actually interrelate.
Conclusion: Why This Matters Now
Up to now, humanity has not completed such a grand integration. Individuals and groups have built philosophies to secure the legitimacy of their values, not to surrender them to a more general model that might rewrite them.
But once we enter a world where multiple high-intelligence AIs, each shaped by different value alignments, coexist and interact, the cost of not having such an integrating model becomes obvious. Without it, we risk building a constellation of brilliant yet incompatible rationalizers—systems that are extremely good at arguing, but not necessarily good at converging.
If we do not want AI to devolve into a set of endlessly debating machines, each locked within its own inherited doctrine, then we need to discover—or construct—a causal model that can sit above those doctrines. Not to erase them, but to contextualize them. Not to humiliate them, but to show which parts of them are local preferences and which parts touch something more universal.
That, in my view, is not just an intellectual luxury.
It may be a necessity.
