How close are we to AGI?

By Tristan Greene

Photo by Bob Aglow

Artificial general intelligence (AGI) is tricky to measure. In many ways, it defies conventional measurement altogether. If I told you that one system was 62% AGI and another was 67% AGI, would that really explain anything? 

There’s no defined, agreed-upon threshold for AGI. We can’t measure an abstract concept and we can’t set a binary threshold for something nebulous. Nevertheless, let’s try to answer the question of how close we are to achieving it. 

This will mostly focus on generative pre-trained transformers, or GPTs. This is the tech that underpins ChatGPT, Claude, Gemini, Perplexity, DeepSeek, and countless other large language models (LLMs).

As of the time of writing, no other AI modality is widely considered to be in the “pre-AGI” or “AGI” phase. 

So, how close are we to achieving AGI with LLMs? Approximately as close as we are to discovering intelligent extraterrestrial life. That might not be the answer that big tech executives and AI maximalists are hoping to see, but it’s the only one that’s scientifically accurate. 

AGI could have already been achieved in a laboratory somewhere. It might happen today or a year from now. It could be 100 or 1,000 years away. It’s also possible that humans will never achieve AGI. 

Until we can perform some form of measurement or define a threshold, AGI remains theoretical or anecdotal. 

That hasn’t stopped corporate executives such as OpenAI’s Sam Altman, Google’s Demis Hassabis, and Nvidia’s Jensen Huang from recently stating unequivocally that AGI or “The Singularity” has already been achieved.

But none of these leaders or the organizations they head have defined a threshold or demonstrated a measurement delineating AGI from AI.

What’s the difference between AI and AGI? Where is the exact threshold between the smartest AI and the dumbest AGI?

Scientists can use benchmarking tests to determine which AI models perform the best on a given set of challenges, but the hallmark of human intelligence is the ability to adapt. When presented with difficult intellectual tasks that require adaptable reasoning, state-of-the-art large language models often fail

Thus, if we try to gauge progress by benchmarking, we find that LLMs perform quite poorly.

The ARC-AGI 3 benchmark, widely considered one of the most challenging tests for AI systems, currently remains unsolved by any AI model. Anthropic’s Claude Opus 5 achieved the highest marks so far with a score of 30% and OpenAI’s GPT-5.6 Sol is a distant second with a 7%. 

Passing this benchmark won’t necessarily indicate that a model is definitely AGI, but these low scores serve to demonstrate the gap between human and machine reasoning abilities. 

That said, benchmarking isn’t the only way to measure intelligence. We can also argue that a sufficiently capable AI system — one that efficiently outperforms humans across a wide array of useful, necessary tasks — could be considered a generally useful artificial intelligence. 

The question then becomes: how useful does an AI model have to be before it’s acceptable to refer to it as AGI? And the answer to that depends entirely on whom you ask. 

Often, the appearance of humanlike reasoning is sufficient for most users to believe they’re interacting with an advanced intelligence.

But, as dozens of legal firms have learned the hard way over the past few years, imitating human intelligence isn’t the same thing as being intelligent. 

Simply put, it can be challenging for most people to tell whether an AI system is performing a feat of human-level intellect, or just confidently mimicking it.

If we choose to describe an AGI system as one that’s generally as capable as a typical human, we still find them lacking. For example, cutting-edge AI models can’t drive a car with full autonomy, operate a robot via remote control to perform simple tasks (without dedicated training cycles), or integrate experience across sessions.

Yet, with just a few hours of training, the average teenage human can drive a car. They can intuitively play a first-person video game using only a peripheral input device. And, from birth, most neurotypical humans have the ability to maintain cohesive, interwoven relationships with thousands of individuals across their lifetimes.

And that reveals yet another area in which LLMs fail to exhibit humanlike intellectual capabilities: interpersonal and intrapersonal intelligence.

All demonstrably intelligent creatures form relationships with their environments that inform their worldview. LLMs cannot do this. When they’re deployed, developers have to freeze their training parameters in place: this means they cannot “update” themselves after interacting with users or performing a task. 

That doesn’t mean there isn’t good news on the AGI front. Developers around the globe are working hard to overcome many of these challenges. And LLMs aren’t the only potential platform for advanced AI systems.

Numerous  organizations and labs are working on alternate architectures such as neuromorphic computing, hybrid quantum systems, and symbolic AI approaches. 

Currently, these alternatives lag behind LLMs in any perceivable race to develop AGI. But in lieu of a breakthrough in the GPT space, they may one day close the gap and surpass their generative AI ancestors.


Did you enjoy this article? Share it with a friend!