Right-Sizing AI: Where Small Language Models Do the Job Better

AI governance

AI Workflow Automation

Blog

An operations team routes every inbound customer email through the most capable AI model the company can access. The model has one job, which is to decide whether each message is a refund request, a delivery question, or a sales lead. It does that job well, the monthly invoice grows with every new customer, and nobody has yet asked whether a question with three possible answers needs the most powerful model on the market.

That question now shapes how mature teams approach AI cost optimization. Most AI work inside a business is routing, tagging and classification, and small language models that work faster and at a fraction of the cost of frontier models. Frontier models stay in the system for the work that genuinely needs deep reasoning.

‍

Why Decision Models Have Entered the AI Conversation

A new category of AI has taken shape within the past three weeks. TypeSafe released Jev on 15 September 2026, a model that returns typed, calibrated decisions in one parallel pass rather than generated text. Cloudflare followed on 1 October with Clef and Clef flash, two decision models with open Apache 2.0 weights hosted on Workers AI, and Cloudflare's own benchmark puts Clef flash at a median decision latency of 38.8 milliseconds.

Early production results are encouraging. In one published evaluation on 98 hand verified production email threads, the production generative classifier reached 47% accuracy, while specialized decision models scored 80% and 82%.

The broader signal for leaders is that AI now works in two modes. Frontier models are best for discovery, system design and novel problems. Narrower models are best for running a proven workflow thousands or millions of times.

‍

What Are Small Language Models?

A small language model (SLM) is a compact language model built or tuned for a focused set of tasks. It runs quickly, costs little per request, and can often be deployed on a single GPU or inside your own environment. IBM and Hugging Face both offer clear primers on how these models are built.

Research from NVIDIA takes the position that small language models are sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems, and that agents invoking multiple different models are the natural choice where general conversational ability is essential.

‍

SLM vs LLM: Which Tasks Need a Frontier Model?

The SLM vs LLM decision is easiest when you sort work by the shape of its output.

Tasks well suited to small language models

  • Routing requests to the right team, queue or workflow
  • Classifying tickets, documents and transactions
  • Extracting fields such as dates, amounts and product codes
  • Policy checks that end in a clear yes or no

Tasks that justify a frontier model

  • Open ended reasoning across many sources
  • Situations the system has never seen before
  • Long form writing, strategy work and complex code

A useful rule follows from this. If the answer comes from a short, known list of options, the task is a strong candidate for a small model. Microsoft's comparison of the two model types reaches a similar conclusion.

‍

‍

How LLM Routing Turns Model Choice Into Architecture

LLM routing is a layer that reads each request and sends it to the model best suited to handle it. The small model resolves the cases it is confident about, and anything uncertain moves up to a frontier model. In the same evaluation, across 31 live inbound emails, the local decision model acted on 8 with zero errors and passed the rest to the frontier model.

This changes AI model selection from a one time vendor decision into a design choice that improves over time. AWS and OpenAI both publish practical guidance on this pattern.

‍

‍

Why Cost per Task Is a Better AI Metric Than Cost per Token

LLM cost is usually quoted per million tokens, but a business pays for outcomes. Cost per task measures the full model spend required to complete one unit of business work, such as one ticket classified or one invoice processed, including retries and escalations.

The difference can be large. TypeSafe's own workflow evaluations put Jev at $0.0004 per case against $0.0304 and $0.0836 for two frontier models, a 76x to 209x spread. These are vendor figures, so each team should measure on its own data, but they show why cost per task is the number that belongs in a leadership dashboard.

‍

A Practical Checklist for Right-Sizing AI

  1. List every AI call in your workflows and record its monthly volume.
  2. Tag each call by output type, either a fixed set of options or open ended.
  3. Build a labeled test set from real business data.
  4. Compare small and frontier models on accuracy and cost per task.
  5. Set confidence thresholds so uncertain cases route to a larger model.
  6. Review accuracy and cost monthly as volumes grow.

‍

What This Means for Data and AI Leaders

The most efficient AI systems are portfolios of models rather than a single model. That moves the work away from choosing the best model and toward designing the routing layer, the evaluation data and the monitoring around it. Each of these is a data engineering discipline that depends on clean labeled examples, reliable logging and well structured pipelines.

At Datum Labs, our AI & ML and GenAI services help teams map their AI workloads, build the evaluation datasets that prove what each model can handle, and design routing layers that send every task to the right model. The result is AI that costs less per outcome and holds its accuracy as volume grows.

The next stage of AI maturity will be measured less by which model a company uses and more by how well it matches each decision to the right level of intelligence. Teams that right-size early will scale AI with confidence.

Heading 1

Heading 2

Heading 3

Heading 4

Heading 5
Heading 6

Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat. Duis aute irure dolor in reprehenderit in voluptate velit esse cillum dolore eu fugiat nulla pariatur.

‍

Block quote

Ordered list

  1. Item 1
  2. Item 2
  3. Item 3

Unordered list

  • Item A
  • Item B
  • Item C

Text link

Bold text

Emphasis

Superscript

Subscript

Featured Insights

October 8, 2026
Datum Labs
Most AI work is routing, tagging and classification. Learn where small language models beat frontier models on cost per task, and how LLM routing works.

Right-Sizing AI: Where Small Language Models Do the Job Better