SLMs: Ethical governors of the future?

By Megan Carnegie

Photo by Bob Aglow

Small language models (SLMs) hold great promise for the future of AI development. They’re faster to train, cheaper to deploy, and easier to work with than large language models (LLMs). SLMs are also emerging as a more sustainable and scalable route to industry AI adoption and scaling advanced AI systems. 

One fledgling use case is that SLMs could act as moral/ethical guardians by monitoring and checking the outputs of larger, more powerful LLMs. The best-case scenario? SLMs could be integrated without major architectural changes to prevent larger models from producing harmful, biased or unethical content — like an ethical governor.

This burgeoning paradigm echoes earlier work on generalist systems such as DeepMind’s Gato, which explored the idea of a single model handling many different tasks. In that sense, SLMs could extend the same general-purpose logic by serving as a safety layer rather than as the main model itself.

So how might that work in practice? One approach is to use transformer-based or instruction-tuned models as the “guardian” component, trained on datasets that equip them to spot risky outputs and enforce safeguards. That could include medical data, as well as datasets designed to test the safety and reliability of language models.

In their framework, AI Researchers at the popular South Korean search engine Naver propose a real-time loop in which the SLM continuously analyses the larger model’s output and intervenes when necessary — for example, by modifying or blocking potentially harmful text — to support safer and more responsible AI development.

Because they’re smaller, easier to inspect, and more tightly constrained, SLMs could offer a way to control more powerful, and potentially riskier, AI systems. But Dr. Matt Hasan, founder of The AI Humanist Movement, says they can only work within a broader governance framework. “SLMs do not create ethics, they only enforce rules that someone else has chosen,” he points out. “If the incentives are wrong, they will simply enforce the wrong things very efficiently.”

However tempting it might be to conjure a quick technical fix for bad actors, SLMs cannot act as the morality police without human oversight, accountability, and a clear line of authority. 

Exclusive: David Gunkel on AI Ethics and the “Relational Turn”

In Part Two of our four-part series, Gunkel explores indigenous kinship ethics, moral standing, and why AI ethics should move beyond harm reduction toward engineering for good.

“If the focus moved away from safety and harm, which are negative kinds of attention, more in the direction of the positive: How do we engineer for good?”

Existing research paints a picture of SLMs as strong contenders to serve as ethical bellwethers. While no research has proved that SLMs could act as a governor in a fully realized sense, there’s plenty to suggest they can be trained to perform governor-like safety functions. 

The Naver team’s paper, for example, finds that a small model can detect harmful queries and generate safeguard responses, sometimes matching or outperforming larger baselines. 

Taking a slightly different tack, evidence from 2025 shows that preference knowledge can be transferred into SLMs in a structured way and improve alignment with human judgments. Smaller models can also inherit chain-of-thought reasoning from larger ones, which would make them more capable of rule-based oversight, according to researchers from the Idap Research Institute and the University of Manchester. NVIDIA researchers also argue that SLMs are well-suited to adapt to changing user needs, output formats, and local regulatory requirements

Taken together, it’s clear that an SLM could sit beside a powerful model as a lightweight filter, a veto mechanism, or a safety response layer.

Can SLMs serve as a complete ethical governance framework, though? The jury is out. Smaller models are no more ethical, by default, than larger ones. They are simply a product of past and present human social realities. As philosopher Shannon Vallor writes in The AI Mirror, all AI, regardless of size, will reflect the errors, biases and failures already embedded in humanity.

Modular systems, with one system that generates outputs and another that checks, corrects or vetoes them, come with their own risks. Dr. Hasan identifies false confidence as the primary failure mode. “Multiple models can agree and still be wrong in exactly the same way,” he says. “Collusion can also happen, especially when models share the same data or objectives.” Other issues include increased latency and greater complexity as systems slow down and humans are bypassed. “There’s diffusion of responsibility: when something fails, no one owns it,” he explains.

Either way, an SLM can still be biased, manipulated or poorly aligned. Examples of constrained and audited LLMs are out there. Anthropic’s Constitutional AI approach trains models against an explicit written “constitution,” with reinforcement learning from AI feedback to produce assistants that are less harmful. In production settings, many teams add guardrail layers that filter inputs and outputs, log decisions and route high-risk cases to humans, while red-teaming exercises probe for bias, toxicity, privacy leakage and jailbreak vulnerabilities before deployment.

In these nascent stages, it’s clear that SLMs can and should contribute to broad ethical frameworks, but there are dangers to relying on them and them alone. Effective governance will demand the separation of powers, with independent models that all perform a distinct role, so that no single system is grading its own homework. “It requires clear escalation paths to humans, full audit logs, and, above all, authority clarity,” says Hasan. “People need to know who can stop a decision, and under what conditions; without that, governance is just safety theater.”

SLMs might well be part of AI governance, but only if society decides to prioritize responsibility, oversight and liability over speed and efficiency. The most hopeful outcome is that they’re tools of enforcement, not substitutions for moral judgment. 

Ethical concerns discussed in this article:

  • How do we define and ensure meaningful ‘human oversight’ of AGI systems at scale?
  • How can we ensure that AGI development is aligned with the broader interests and values of global society, rather than those of a narrow group?
  • Who is ultimately responsible/culpable for the actions of a given AGI?

Get sharper analysis on AGI ethics

Original essays, expert commentary, and curated analysis on the human stakes of AGI.

Subscribe to AGI Ethics News


Did you enjoy this article? Share it with a friend!