Return of the small language models

By Tristan Greene

Photo by Bob Aglow

Hello to our new subscribers and a hearty welcome back to our returning readers! The past month has seen substantial growth for AGI Ethics News on social (make sure you’re following us on LinkedIn) and in our behind-the-scenes projects such as the Spark Hunter Survey.

Not to spoil the surprise, but we conducted a deep dive survey to gain insight into how science and technology journalists view the field of artificial general intelligence. And the data I’ve gleaned so far is fascinating.  

We’ll be revealing the anonymized results of that survey along with detailed expert analysis in next month’s issue and unique perspectives on the future of AGI and journalism from some of the world’s top writers.

You can also look forward to part three of our interview with philosopher David Gunkel (read part two in this month’s edition).

For this April edition, we’ve chosen the topic of small language models (SLMs). These are a facet of AI development that’s largely been forgotten in the conversation about advanced artificial intelligence. 

SLMs are often confused for AI agents or thought of as one-trick ponies. But it’s more useful to think of them like “focused” language models than single-domain “experts.” 

New research demonstrates that SLMs can rival large language models (LLMs) across a variety of tasks. They may even hold the key to developing safer, more ethical AI systems

In a recent study, researchers from San Francisco State University, Carnegie Mellon, Columbia, Cornell, and Pennsylvania State found that SLMs were consistently more efficient than LLMs:

“Small models (0.5 – 3B) achieve superior Performance-Efficiency Ratios (PER) over larger counterparts, which exhibit diminishing returns beyond 14B parameters. While medium-scale models (3 – 14B) offer an optimal performance balance, 0.5 – 3B models are most effective for resource-constrained deployments. These findings provide a quantitative basis for task-specific selection, challenging the ‘bigger is better’ paradigm.”

The numbers above represent parameter size displayed in billions, a general measure of a transformer model’s size. For comparison, experts believe GPT-5 has somewhere between 1.7 trillion and more than 10 trillion parameters. 

Here’s how SLMs work

Small language models are, simply put, natural language processing models that are smaller in scope and scale than large language models. There’s no defined universally agreed-upon “cut off number” or size requirement to differentiate the two because SLMs are either intentionally trained on curated, high-quality data or created as derivatives of LLMs. 

In practice, LLMs are “spun up” (programmed, trained, and encoded for use as “foundational models”). From there, developers will then create fine-tuned, bespoke small language models by compressing, distilling, and quantizing LLMs. 

The process is a bit more complex than the above paragraphs would make it seem but, ultimately, the goal of most SLMs is to create a more efficient, streamlined AI model that’s tuned to better accomplish a specific set of tasks. 

Some of the most popular SLMs, such as Mistral’s 7B and 3B models, boast comparable benchmarks across specific domains to LLMs while operating at much greater efficiency. 

A key advantage, for many SLMs, is that they can be run using CPUs and RAM. This means they can, ostensibly, be operated on smartphones, laptops, and just about any modern equivalent computing device. This is in juxtaposition to the GPU farms and data centers required to operate many LLMs. 

Because SLMs can operate on-device, they provide myriad data and network security advantages over LLMs accessed via the cloud.

As Megan Carnegie discusses in her return article for this month’s issue, one of the benefits to having a smaller footprint is that SLMs have the capacity to serve as “governors” in AI systems that combine SLMs with LLMs. 

In this paradigm, you might fine-tune an SLM to read the outputs generated by an LLM and ensure they’re in line with the prescribed ethical parameters. This would, hypothetically, create a framework wherein the AI model is free to explore its own ethical answers — perhaps for developmental or research purposes — while the SLM ensures end-users aren’t exposed to unwanted outputs such as private or harmful data. 

Thomas Macaulay also returns this month to weigh in on SLMs with an article discussing how they could function as ethical monitors and controllers overseeing swarms of AGI agents.

Furthermore, this issue features an analysis of a pre-print from Princeton titled “An Alternative Trajectory for Generative AI,” which could have immediate implications for both small language models and the future of ethical AGI.

Thanks for subscribing, and please feel free to drop us a line at Editor@AGIEthicsNews.com anytime!

Get sharper analysis on AGI ethics

Original essays, expert commentary, and curated analysis on the human stakes of AGI.

Subscribe to AGI Ethics News


 

Did you enjoy this article? Share it with a friend!