Endgame: How far can large language models take us?

By Tristan Greene

Photo by Bob Aglow

There can be no thorough discussion on the goal of creating artificial general intelligence (AGI) without mentioning the elephant in the room: large language models (LLMs).

People often conflate LLMs with all artificial intelligence. Most folks don’t consider the algorithms powering their cellular networks and the transportation industry when they ponder AI. They’re thinking about ChatGPT, Claude, Gemini, and the other chatbots they interact with.

If you asked 100 people if they thought the Netflix app is conscious, sentient, or sapient, you probably wouldn’t get any “yes” responses. Ask the same people whether ChatGPT or Claude is and it’s almost certain you would.

Does this mean that LLMs are more powerful, capable, or humanlike than the algorithms that power Netflix’s recommendation system? No. They operate using the same fundamental computer processes. Arguably, the only difference is how the end user perceives the interaction. 

LLMs are difficult to observe objectively because, unlike other nonhuman systems such as climate, they appear capable of interacting with us directly. And they seem to communicate with us in real time using our own language.

We can argue that, for example, the climate speaks directly to us. But it usually takes a while for it to respond. We’d probably feel different about how we’ve impacted our environment if lightning struck nearby and thunder boomed every time we littered or dumped trash in the ocean. 

This general subjectivity toward LLMs doesn’t just affect laypersons. Even researchers at the cutting-edge of computer science and artificial intelligence development disagree on how humanlike LLMs are or will become.

“The problem that makes this particular situation awkward is that you cannot really calibrate what it means to have human-like attributes in anything that is related to an LLM. If you have a thermometer and you’re doing an experiment with temperature, and you want to verify that your experiment worked or not, you can calibrate your tool. You cannot do that in an LLM.”

As Adrian de Wynter, principal applied scientist at Microsoft, wrote in a recent preprint paper published to arXiv, “even though large language models (LLMs) are comparatively new, they are both widespread and poorly understood.”

AGI Ethics News spoke to de Wynter in a recent interview to get his thoughts on the goal of developing advanced AI systems such as AGI and why even the experts can’t seem to agree on what, exactly, LLMs are. 

According to de Wynter, there’s more at play than just the simple interaction between a human and a chatbot. “The problem that makes this particular situation awkward is that you cannot really calibrate what it means to have human-like attributes in anything that is related to an LLM. If you have a thermometer and you’re doing an experiment with temperature, and you want to verify that your experiment worked or not, you can calibrate your tool. You cannot do that in an LLM.”

Basically, under de Wynter’s premise, if a researcher assumes that LLMs do or don’t have humanlike traits, that assumption will color their work. 

In de Wynter’s recent paper, titled “If LLMs Have Human-Like Attributes, Then So Does Age of Empires II,” he shows that LLMs are non-unique and illustrates this by building a working, Turing-complete neural network inside the 1999 strategy video game Age of Empires II (henceforth: AoE2).

Essentially, de Wynter built circuits inside the game by turning a massive map into a compute space and using in-game assets as switches inside of functionally-mapped circuits. The result was a version of an LLM that operates the same, except the interface. Rather than using electronic switches inside of a GPU cluster, users see digital goats and villagers walking in constrained paths to generate outputs.

The prompts used and the outputs generated remained consistent between the regular version of the LLM and the goat-powered one, but the user experience was vastly changed.

If we think about ChatGPT, Claude, Gemini, etc., in a void, it’s easy to view them as humanlike. We talk, some electronic magic happens inside of a black box, and the machine talks back. 

LLMs also present as technologically special. We know it takes billions of dollars to train and operate them. Some of the world’s best scientists are collaborating across disciplines to develop bigger and better LLMs. And about 1/8th of the world’s population uses one on a weekly basis. They’re obviously a big deal. 

But, as de Wynter’s work demonstrates, LLMs aren’t unique. The same computational processes used to generate LLM outputs via switches in GPUs (and goats and villagers in AoE2) can also be demonstrated using any similar substrate such as “the Greater Boston area,” or circuits built out of LEGO. 

“Given that LLMs are sufficiently effective at mimicking anthropomorphic attributes to some extent,” he writes in the paper, “it follows that said mimicry — or true anthropomorphic behaviour, depending on the view — is not specific to LLMs as entities existing within a computer.” 

This begs the question: is there a computer architecture capable of producing an intelligence that appears to possess humanlike consciousness, sentience, or sapience, that fails to maintain the humanlike state under alternate substrates?

In layperson’s terms: If LLMs appear equally humanlike whether they’re running inside of a video game or producing outputs via an analog computer made of LEGO, then it’s hard to argue that anything “special” is going on when we energize the model. 

If we view LLMs as stepping stones toward a new modality, or the recipe for a future machine that can retain its inalienable “selfness,” then the best path forward might be to exhaust the limits of the LLM architecture in the same way we’re attempting to do in the semiconductor industry with regard to Moore’s Law

Thus, if the goal is to develop the most convincingly humanlike architecture, it’s clear that LLMs are the way forward in the short term. But, if our goal is to develop machines that are scientifically unique, one of the base precepts underpinning humanity’s view of itself as humanlike, then perhaps LLMs are a mere diversion — whether they’re powered by Nvidia’s GPUs or AoE2’s goats and villagers. 

Ethical concerns addressed in this article:

  • What guidelines address the anthropomorphization of AGI and its effects on perceptions of agency and moral standing?
  • What safeguards ensure an AGI’s subjective experiences are not simulated but genuinely arise from its architecture?
  • What risks arise from AGIs that may have “pseudo-consciousness,” mimicking the building blocks but lacking genuine subjective experience?

Get sharper analysis on AGI ethics

Original essays, expert commentary, and curated analysis on the human stakes of AGI.

Subscribe to AGI Ethics News


Did you enjoy this article? Share it with a friend!