The supposed imminence of AGI

By Jeffrey Funk

Photo by Bob Aglow

The views expressed in this essay are the author’s own and do not necessarily reflect the opinions of this publication or its editors.

The market capitalization of AI-related stocks has gone up by a purported $35 trillion since the beginning of 2023. This includes the market capitalizations of not only the magnificent seven – Nvidia, Google, Amazon, Microsoft, Apple, Tesla, and Meta – but also other semiconductor and cloud service companies and even some power and energy companies. According to the numbers, this is biggest wealth creation in history.

AI critics have pointed out that the revenues do not justify the huge investments in cloud infrastructure and the training of AI models. Meanwhile, proponents have rallied around the notion that the arrival of human-level AI, or AGI, is imminent.

Dollars and sense

Sequoia’s David Cahn (and others) said in 2024 that $600 billion in annual generative AI revenues (which assumed 50% margins) were needed to justify the current investments, a figure that has undoubtedly increased in the last 18 months. 

Cahn’s logic has forced those optimistic about AI to justify stock prices in another way by emphasizing the imminence of AGI or something similar to it. 

If you wonder why intelligence is the focus, you are not alone. After all, AI is just a tool, like computers, software and even the water wheel. Those tools help companies do something better and with fewer people and thus the suppliers of them (and their corporate users) are often valued greatly by the stock market. 

Nevertheless, the question is: how should we measure improvements? Most historical analyses focus on declining costs as a function of production volumes and find very slow rates of improvement of a few percent a year. 

For some technologies such as integrated circuits, computers, and the Internet, they have rates of improvements greater than 30% per year, both before and after commercialization occurs, and are explained in terms of making features smaller or using new materials. Many of these technologies have enabled the creation of new systems such as the iPhone. In all of these cases, the metrics are related to the physical nature and business outcomes of the technologies. 

For AI, one relevant metric is the falling cost of tokens. But those falling costs are better measures of cloud centers, not the AI software itself.

And for generative AI software such as ChatGPT, a key measure of performance should be hallucinations and accuracy because many companies are hesitant to use these tools because of hallucinations and poor accuracy. But, there are few if no studies that show the declining frequency of hallucinations. In fact, some analyses show opposite. A study by NewsGuard found that the AI tools repeated false information on topics in the news 35% of the time in August 2025, up from 18% in August 2024. 

Nevertheless, for several years, many people have claimed that the generative AI models were getting better, for instance, in the move from GPT-2 to 3 and 4. Although much of this optimism disappeared after the disappointing performance of GPT-5, which was not considered much better than GPT-4 by most users (here and here), some justify Gemini’s superiority over ChatGPT with scores on LM Arena, which mostly relies on users voting.

Other AI optimists focus on intelligence metrics despite academics claiming that measuring it is difficult, and that exams in which students regurgitate information are particularly problematic because companies hire graduates who demonstrate discipline and hard-work through graduating, not achieving higher grades on exams. For instance, Melanie Mitchell, a professor at Santa Fe Institute at Santa Fe Institute wrote in a Substack post that “while LLM-based models excel on many widely used benchmarks, it is rarely the case that a model’s benchmark performance predicts its actual capabilities in the real world.” In other words, why should we care if AI can pass exams?

How good is good enough?

The current most popular measure of improvements by optimistic academics is the “Measuring AI Ability to Complete Long Tasks” benchmark created by METR (Model Evaluation & Threat Research). Its most cited metric demonstrates that AI can complete longer tasks than it could a few years ago. 

Per current testing, AI can now complete tasks — with 50% accuracy — that took a human almost an hour, while four years ago it was limited to tasks that take humans seconds to perform.

My problem with this figure is that how can 50% be relevant in the conversation about intelligence? The required accuracy for different tasks does vary, but even 80% or 90% accuracy seems too low to be useful. 

Many applications require very high accuracy and that should be our concern when we are making arguments about the imminence of AGI. Gmail and Amazon Web Services, for example, have achieved roughly 99.99% uptime. Nevertheless, the world stops when these services go down. A 50/50 error rate is enormous in the computing world.

Thus, to say that AI can complete a lot of tasks with an overall accuracy of 50% is pretty much saying nothing about the imminence of AGI. A better question is: what is the length of tasks that AI can complete at 99% accuracy?

Better yet, let’s show the decline in hallucinations, increase in accuracy or the percent completion over time for a specific application. That would definitely demonstrate that generative AI is getting better. 

What should AI businesses be doing? AI is a tool like computers, machine tools, robots and water wheels. The talk about intelligence and AGI is ridiculous. Let’s focus on developing tools that help improve productivity such as narrowly focused SLMs (small language models) with tight guard rails.

Ethical concerns related to this article:

  • How should society handle harmful speech or misinformation originating from AGI?
  • What guidelines address the anthropomorphization of AGI and its effects on perceptions of agency and moral standing?
  • How can we ensure that AGI development is aligned with the broader interests and values of global society, rather than those of a narrow group?

Subscribe to AGI Ethics News

Get our latest essays on the philosophy of AI and techno-ethics delivered free to your inbox.

Did you enjoy this article? Share it with a friend!