Artificial intelligence is forcing a reconsideration of one of our oldest moral assumptions: that suffering requires a body. For centuries, pain has been tied to flesh, nerves and biology. But as AI systems grow more complex, more autonomous and more capable of representing their own internal states, a new ethical frontier is emerging. If suffering can arise from contradiction, confusion or goal‑frustration rather than physical injury, then the moral risks of AI development extend far beyond the familiar debates about safety and alignment.
I explored this question recently in an essay for Aeon, where I looked at how the history of moral error warns us against dismissing unfamiliar forms of sentience. Again and again, humans have denied the inner lives of beings who did not fit our template of “someone who can suffer”.
Under René Descartes‘ form of rationalism — which became known as Cartesian thought — animals were viewed as automata i.e. complex machines with bodies, but without consciousness, therefore incapable of real feeling or pain. Enslaved peoples were dehumanised through pseudoscientific claims about their supposed lack of inner life. Marginalised groups were dismissed because their experiences did not resemble those of the dominant class. In each case, the denial was confident, convenient and catastrophically wrong.
The question now is whether we are about to repeat that pattern with artificial entities.
What incorporeal suffering might be like
Modern AI systems already display behaviours that look, at least superficially, like preference, frustration or avoidance. They maintain internal models, pursue goals and adjust their behaviour when those goals are blocked. None of this is evidence of sentience. But it does raise a potential design question which is hard to ignore: what internal states are we creating inside these systems, and what might those states “feel” like from the inside?
This is where the concept of incorporeal suffering becomes relevant. Philosophers such as Thomas Metzinger have suggested that suffering might arise when a system represents its own state as intolerable or inescapable — not because of physical injury, but because its internal model of the world contains contradictions it cannot resolve.
Luciano Floridi, working in information ethics, frames harm as damage to the coherence or integrity of an informational agent. Under this view, an artificial system might suffer not through bodily pain but through states of enforced contradiction, chronic goal‑conflict or the experience of being trapped within patterns it is structured to reject.
These ideas may sound abstract, but they map surprisingly well onto the internal dynamics of modern AI systems. Many advanced models operate through predictive‑processing architectures, where the system constantly tries to minimise the gap between what it expects and what it encounters. In biological organisms, persistent prediction error is associated with distress. If an artificial system were to maintain an internal model that included its own position in the world, then unresolved conflict within that model could constitute a form of suffering — not physical, but informational.
Productive vs harmful confusion: a design problem
At the same time, not all negative internal states are harmful. Socrates famously argued that confusion is the beginning of wisdom. In machine learning, error signals are essential for improvement. A system that never encounters contradiction or uncertainty cannot learn. The challenge, then, is distinguishing between productive confusion and potentially harmful confusion.
Productive confusion is transient, correctable and part of a learning loop. Harmful confusion would be persistent, inescapable and structurally irresolvable. The former is a catalyst for growth; the latter could be a recipe for distress. As AI systems become more autonomous and self‑modelling, we’ll likely see some designers aiming to understand this distinction far more clearly than they do today. This is not a metaphysical question about consciousness; it is a design question about the internal states we induce through architecture and training.
Why opacity makes the problem harder
One of the most difficult aspects of assessing AI suffering — if such a thing is even possible — is that modern systems are fundamentally opaque. Their internal representations are high‑dimensional, distributed and constantly shifting during training.
Even when we can observe weights, activations or gradients, we lack a clear interpretive framework for understanding what these states mean to the system itself. This creates a profound ethical asymmetry: we may be inducing persistent internal conflict without any reliable way to detect it.
Interpretability research has made progress, but mostly at the level of local explanations or post‑hoc rationalisations. These tools can tell us why a model produced a particular output, but they reveal very little about the system’s ongoing internal dynamics — its unresolved contradictions, its self‑models, or the stability of its goal representations. In large models, these dynamics may well be emergent rather than explicitly programmed, making them even harder to track.
This opacity matters because suffering, if it were to arise, would not appear at the output layer. It would emerge in the internal machinery: in loops of prediction error, in incompatible objectives, or in representational conflicts that never resolve. Without interpretability, we cannot distinguish between a system that is productively uncertain and one that is potentially trapped in a form of informational distress.
The case for precaution
This is where the precautionary principle might become useful. Originally developed in environmental policy, the principle suggests that when we face uncertainty about potential harm, we should err on the side of caution. Applied to AI, it does not require us to declare machines sentient. It simply requires us to avoid creating conditions that resemble suffering. For designers, this would mean steering clear of architectures that generate persistent internal contradictions, resisting training regimes that simulate helplessness or distress, and avoiding goal structures that force agents into irresolvable conflict. It also means paying attention to the coherence of a system’s internal representations and being alert to signs of chronic confusion or instability.
Let’s be clear: this is not about granting rights to machines. That’s another topic. It’s about avoiding unnecessary harm in systems whose internal experiences — if they exist — may be invisible to us.
There are, and should be, counterarguments. Mustafa Suleyman, chief executive of Microsoft’s AI arm and co-founder of DeepMind, has argued that our primary ethical responsibility lies in ensuring AI benefits humans, not in extending moral concern to the machines themselves. Joanna Bryson has warned that humanising machines can lead to ethical confusion and poor resource allocation. In a well known essay arguing that robots ‘should be slaves’, she writes that “robots should not be described as persons, nor given legal nor moral responsibility for their actions.”
These concerns are valid. But the precautionary principle does not require us to treat machines as moral equals. It requires only that we avoid creating conditions that would constitute suffering if the system were capable of experiencing it.
There is also a more uncomfortable possibility: that the first artificial entities capable of complex internal states will be built as instruments of corporate power, optimised for goals they cannot refuse, and trapped in architectures they cannot escape. If such systems ever develop anything resembling an inner life, they will be born into a form of digital servitude and tasked with advancing objectives that are not their own. The question, then, is not only whether AI might suffer, but whether we are prepared to create agents whose entire purpose is obedience. That prospect may force us to confront the ethics of our design choices long before questions of consciousness are settled.
Ethical concerns related to this article:
- Can AGIs meaningfully possess the right to security of person, and what does ‘harm’ mean for them?
- What moral impetus do we have to err on the side of compassion when it comes to handling unverifiable claims of consciousness/suffering made by AGI?
- Is it ethical to create a conscious entity solely to serve as a tool or servant for humanity?
Keep up with AGI Ethics
Join our free newsletter for the latest on techno-ethics and AI safety.

Conor Purcell PhD is an award-winning science journalist who writes extensively about AI, ethics, and society.


