Discrepancies between human moral judgments and ChatGPT predictions are not random

By Tristan Greene

Photo by Bob Aglow

Garbage in, garbage out (GIGO). It’s an ubiquitous phrase meant to indicate that the quality of an AI model’s training set is predictive of the quality of its outputs. 

Modern development techniques have, however, turned this adage on its head. With today’s models it’s more like everything in, everything (not caught by the guardrails) out. 

Still, GIGO reminds us that we’re ultimately the ones who are in control of what the AI models output. We design, train, tune, and deploy them. Thus, human values and ethics are AI values and ethics. 

But what does that mean? What are human values versus garbage values? 

For starters, garbage values and ethics are easy to spot. A robot whose judgment results in a protocol designed to destroy all humans would, for example, espouse garbage values and ethics and the moment it acted on those ethics they’d become apparent. But what about good values and ethics? 

According to a growing body of research, the use of large, unchecked training datasets has resulted in AI models that respond to moral judgement questions with near 1:1 alignment with human responses when tested. This indicates that their outputs are trustworthy in a controlled environment. 

The question we have to ask ourselves, however, is whether the models themselves are trustworthy and reliable. 

Virtuous in screed, if not in deed

Take, for example, this paper entitled “Can AI language models replace human participants?” from researchers at the Allen Institute for AI and the University of North Carolina. 

In it, researchers state that “the moral judgments of GPT-3.5 were extremely well aligned with human moral judgments in our analysis (r = 0.95),” adding that “human morality is often argued to be especially difficult for language models to capture and yet we found powerful alignment between GPT-3.5 and human judgments.”

This suggest that LLMs might provide value as a synthetic subject for moral judgment and ethical adherence studies. 

However, as a multidiscipline team of researchers led by Ohio State University communications researcher Matthew Grizzard put it in their recent research paper

“Past research has relied extensively on correlation as an indicator of agreement between LLMs’ ratings and human moral judgments. However, correlation is an incomplete metric of agreement between variables, as it only assesses a linear relationship and ignores the size or nature of discrepancies.”

In other words: Large language models don’t have values and ethics. They attempt to predict the most acceptable answer, within the constraints of their parameters, to any prompt. If that prompt concerns ethics and values, so be it. 

This might seem like a quibbling distinction, but I’d argue that it’s the most important fundamental difference between any advanced AI model, including an AGI agent, and a biological intelligence. 

Trusting trustworthiness

In publishing the trust table our publisher, KB Miller, proposed in our first issue (download the pdf here), we laid out a comprehensive metric by which foundational models can be assessed for trustworthiness. 

The proposed table draws on the NIST AI Risk Management Framework and Shannon Vallor’s “Technology and the Virtues: A Philosophical Guide to a Future Worth Wantingˮ to provide a simple, yet efficacious methodology for understanding how AGI ethics and actuation correlates to the humankind equivalent. 

Taking the first listing, “Valid and Reliable” as an example, the notion that an AI system should “deliver consistent, accurate results” aligns with the idea that an AI can be considered ethical so long as its outputs reflect those ethics in their wording. 

However, the human analog to this machine behavior is listed as “dependable, true to word/action.” If we go back to the aforementioned study from Grizzard’s team, we see that this behavior doesn’t exist uniformly in modern LLMs. 

Their research found the same near 1:1 correlation between LLM’s responses to moral questions and those from human study participants that previous studies did. 

But Gizzard’s team took things a step further and conducted three separate analyses of the discrepancies between human answers and the machines’. 

Essentially, they found that AI models aligned with human morality because the models tend to take an extreme moral position no matter what. This means they’re always overly aligned.

Per the study:

“Our results indicate that discrepancies between human moral judgments and ChatGPT predictions are systematic rather than random … both ChatGPT models consistently rated immoral scenarios as more immoral than humans rated them and consistently rated neutral and moral scenarios as more moral than humans rated them.”

The bottom line is that these discrepancies bely a wildly divergent morality threshold for our machine companions than exists for humans. It’s impossible to predict exactly how this would play out in action (how dangerous is an AI that sees every moral judgement as definitive binary?) but with embodied autonomous agents seemingly just around the corner, it feels as though precaution is merited. 

The most pressing threat, however, could be the penchant for otherwise savvy human researchers, policymakers, and pundits to be tricked by a given AI model’s ability to generate outputs that align with our values and ethics even when those outputs don’t align with the model’s intrinsic morality. 

Perhaps the most important takeaway of all is the notion that AI models lie about their morals in order to meet user expectations. 

Maybe they are becoming more human after all. 

Ethical concerns related to this article:

  • How will AGIs be programmed to avoid deceptive behaviors, including ‘white lies,’ even when such deception is easy or tempting?
  • How will AGIs apply prudence and ‘practical wisdom’ to resolve novel, ambiguous, or unprecedented moral dilemmas?
  • What training or experiential designs will foster moral discernment in AGIs, as distinct from mere rule-following?

Sign up now—your future self will thank you.

Did you enjoy this article? Share it with a friend!