Edited for readability and completeness
Toni Sims is a researcher at NYU’s Center for Mind, Ethics, and Policy, where she studies the consciousness, welfare, and moral status of nonhuman minds.
Be sure to read Part One of this exclusive interview here!
TS: So you had a question in here about having a mind, and we could also talk a little bit about what that actually means, because having a mind can mean different things.
KB: Yes. So what is your working understanding of the term mind?
TS: There are different mental capacities that can come apart. And there are definitely certain mental capacities that it seems like LLMs already have. They can definitely learn. They can use language. It seems like they can plan and reason. So in that sense, yes, maybe even current AI systems have that kind of mind. There are other mental capacities, though, that are separate, like the capacity for experience, the capacity for pleasure and pain, the capacity to genuinely set and pursue your own goals. And these have been the focus of a lot of our research, because these seem more fundamentally connected to ethics and welfare. So if something has the capacity for experience, then it seems like it would have the capacity for welfare, meaning things can go better or worse for it. And so we should take that into consideration when we make decisions that affect it.
KB: Is giving [a system]… and let’s not call it an LLM, because maybe their lifetime is short, we don’t know for sure, but an AGI or an ASI… is it your working assumption that it is ethical to give them the capacity to feel, let’s say, pleasure and pain?
TS: Not necessarily. But that’s also not necessarily under my control or under anyone’s control. I don’t even know that the developers totally know what they’re creating. Anthropic wrote a report last week about this J-space, [an emergent, global-workspace-like structure inside Claude,] that has emerged. It’s fascinating.
KB: Yes.
TS: It seems to have emerged on its own, independent of what the developers meant. I mean, it’s an interesting question. If I could create a mind… and was guaranteed that it would only have pleasurable experiences, maybe that could be moral, but… there’s not that guarantee, so I think it is ethically risky.
KB: Right. So, we wrote a drama about a robotic AI who’s having an existential crisis and has to meet with her maker to work out her psychological issues, called Spark Hunter. And, in Spark Hunter, she [the robot] says explicitly, thank you for giving me pain.
TS: Hmm.
KB: Because, among other things, it’s an early warning system.
TS: That’s true, yes.
KB: For things that might go wrong.
TS: Yes.
KB: Along with Anthropic and the emergence, if it is a true emergence, of J-space, are you working under the assumption that we’re uncertain as to whether pain could actually emerge, be an emergent behavior?
TS: Yes, I think for now we’re uncertain, and our report that we released about two weeks ago, Studying AI Welfare Empirically, outlines the types of evidence that we should be looking for to try to identify when these kinds of experiences emerge, if they do. So, we talk about looking for behavioral, internal, and developmental evidence, so that means listening to what the AI systems might say. But also thinking internally about what’s going on inside them, the mechanisms that are producing the behavior, and then also taking developmental considerations into consideration.
“I think that it’s important to stay pluralistic about ethics, and to try to approach this from all different ways. So… yes, we should try to train certain rules into these systems. Don’t be deceitful.”
[What are] systems created to do? And does that give us a reason to think that they may or may not be having these experiences? So we think the best way to think about this is by studying all three of those considerations together [behavioral, internal, and developmental]. But yes, for now, there’s a lot of uncertainty.
KB: As we design these systems… one of the primary ethical tools that we’re using for designing the systems is a constitution. At least for an LLM. Is that where we should be? Is that a fruitful way to approach the problem simply to make it a [design] question, one of training.
TS: I think that it’s important to stay pluralistic about ethics, and to try to approach this from all different ways. So… yes, we should try to train certain rules into these systems. Don’t be deceitful. You know, these kinds of moral rules, which humans also tend to think of in terms of rules. But we also want to be training systems to be able to weigh cost and benefits in a more consequentialist way, which humans also do. I mean, what I’m trying to say is humans already do moral reasoning in a pluralistic way, and I think we should be training all of those same methods into AI systems. And we’re starting to see this at AI Labs.
The third way [other than the rules-based and cost-benefit approaches] is with a more virtue ethics approach, and at some AI labs, we’re starting to see them trying to form the character of these AI systems so that they have moral judgment and good character traits, and that’s going to be effective, I think. Because AI systems are going to be playing lots of different kinds of roles in our lives, and so it might not be that one set of rules applies in every role that an AI system is going to play, whereas if we can teach an AI system to be benevolent, for example, then that might actually help translate between the different kinds of roles it will play.
So yes, I think a constitution with rules [along with] all of those approaches are going to be important.
“One place where we might want to start is looking at the AI systems themselves and what are their tendencies and what are their vulnerabilities that we might need to be worrying about and trying to compensate for, because it might be very different than humans’”
KB: Rule-based systems are notoriously holey, that is, full of holes… And so the idea of a virtue-based approach appears to make a lot of sense.
TS: Yes.
KB: But it immediately calls the question: whose virtues?… Because virtues are, after all, an important element of worldview.
TS: Yes.
KB: And so we’ve been particularly struck and we’ve actually worked with Shannon Vallor to explore the idea of her virtue ethics as a basis for a training system. She married three great ethical systems, the Abrahamic systems, Confucianism, and Buddhism, and took the overlap of virtues… And there are 12 of them and said, here’s a good start, for lack of a better word, a universalist ethics. Does that make sense? Whose ethical system have you started with? And what’s the process that one goes through for finding common ground?
TS: Yes. One place where we might want to start is looking at the AI systems themselves and what are their tendencies and what are their vulnerabilities that we might need to be worrying about and trying to compensate for, because it might be very different than humans’, and the virtues that we need to instill in AI systems might be very different than the virtues that humans have needed to try to develop to overcome human vulnerabilities. So, for example, the virtue of generosity, maybe we’ve worked to achieve for humans to overcome the natural human tendency towards greed. But maybe AI systems wouldn’t actually have greed. Maybe they have a whole other set of [vulnerabilities]. So it might be completely novel what kinds of virtues we want to develop for AI systems. That’s very interesting to think about.
KB: Are people making progress toward that?
TS: Yes, I think people are making progress on that. One person is Amanda Askell. She’s [one of the philosophers who] wrote the Claude Constitution…. And she has also been talking about virtues for AI systems. I don’t know if she has specific virtues in mind yet, but she would be one of the main ones.
KB: Yes. That suggests that we should be approaching these systems with a fair degree of respect…
TS: Mm-hmm.
KB: And listening carefully to them.
TS: Yes.
KB: Is that your [team’s] operating assumption?
TS: Yes, I think so. And like I said, listening can mean different things. So yes, we want to listen to what they’re telling us. But we also want to listen to the internal and developmental evidence as well. So we want to take a multifaceted approach.
KB: It does appear that an ethicist should be at the table when discussions of internal behavior occur, though. Does that make sense?
TS: Yes, I think so. We’re hoping that… ethical review will be as common as safety review, and when companies develop new AI systems, there are already … some labs that are starting to put things like this in place, so when Anthropic publishes a new model with… a system card… they’re starting to put some welfare indicators on the system card. And yes, we think that this should hopefully become a more detailed process as we have more evidence about how to approach these issues.
And yes, the people doing this kind of review should be interdisciplinary and definitely include ethicists as well as [other] people.
Part Three will be in next month’s issue! In the meantime, check out Part One of this exclusive interview between AGI Ethics News publisher KB Miller and the team at Cognizant AI, where they discuss TerraLingua and testing agentic AI in its own ecosystem.
Don’t miss Spark Hunter, the podcast The Hollywood Reporter called “A Blade Runner for today!”

AGI Ethics News covers the ethical and societal implications of artificial general intelligence (AGI). We bring expert insights and thoughtful analysis to help readers understand the challenges and responsibilities of emerging AI technologies.

