Skip to content
Brian Sithu
Go back

The smallest model in your agent stack

Cover image for The smallest model in your agent stack

Decision models answer one question: what should the agent do next. Jev named the category, AWS and Cloudflare shipped open follow-ons, and I would rebuild my agent loop around one.

Brian Sithu

Written by

Brian Sithu

View profile

On this page

03 sections

I spent this year making my agents smarter by giving them bigger models. The last month convinced me I had it backwards for about half of what an agent actually does.

On September 15, a startup called Typesafe AI launched Jev, the first of what it calls System One models. Jev does not write text. You hand it a question with a fixed set of allowed answers, and it returns typed probabilities over those answers, calibrated so that 80 percent means something close to 80 percent. Typesafe calls this a decision model, and with Jev it introduced the category. Classifiers existed long before, so I would not call it the first model of its kind. But it named the idea, and the naming mattered, because two weeks later everyone else started building them.

Here is the job it does. An agent run is full of small closed questions. Which route does this request take. Which tool should fire. Is this action safe to run unattended. Should a human look at this first. I used to spend frontier model calls on all of those, then parse sentences like “I think the best tool here is…” out of generated text and hope the format held. A decision model answers the question directly, as numbers.

The split I now design around: a small model decides at every gate, the frontier model only does the work that clears

Then October 1 brought two more, on the same day. AWS Strands Labs open-sourced Strands Decider 2B: a Qwen3.5-2B with its language model head removed and a small pointer head in its place, released under Apache-2.0 with weights and training scripts included. AWS measured it answering in around 115 milliseconds on an RTX 3090 and 153 on an M3 MacBook. Its announcement post says tens of milliseconds, which reads like the small-task end of its own range, so I trust the measured figures over the slogan.

The same day, Cloudflare shipped Clef and Clef-flash on Workers AI, Qwen-based models with open weights under Apache-2.0. Cloudflare reports median latency of 209 milliseconds for Clef and 39 for Clef-flash, against 524 for Jev in its own evals. Typesafe, meanwhile, claims 70 to 500 milliseconds for Jev. Those two accounts disagree about Jev by enough that somebody’s benchmark flatters its author. Both numbers come from the vendors, so treat them that way until someone independent measures.

Where a 2B model wins

Latency budget is the one I feel first. An agent loop pays for every decision serially. Shave two seconds off forty gates and you gave back over a minute per run. A hundred millisecond gate disappears into the noise. A two second one is the run.

Calibration is the one I trust most. A frontier model asked whether an action should run gives you a sentence. You wanted a number you can set policy on: above 95 percent it runs, below 60 percent it escalates, between those it gathers one more signal. Decision models return probabilities trained to mean what they say. Text answers make you infer confidence from tone. Numbers let you write rules.

Retry cost is the one that compounds. Most of what a gate does is say no cheaply. Checking which tool fits with a 2B model before firing a big generation means the mistakes cost milliseconds instead of full retries. Put the cheap model where the volume is.

What changes at the gate: a sentence you have to parse versus probabilities you can set policy on

Where it loses

Anywhere the question was not enumerated in advance. A decision model answers from a fixed set. It cannot notice the situation your list missed, cannot exercise judgment on something novel, cannot write the response or plan the workaround. The open world stays with the big model. My mental model is simple: the small model runs the gates, the big model does everything the gates let through.

Rebuilding my loop

My loop today looks like every loop. Observe, ask the model what to do, do it, repeat, with the same frontier model deciding and acting at every step. The rewrite keeps that shape but welds the small model to each gate. The decision model picks the route, picks the tool, scores whether the action is safe. The big model wakes up to execute what cleared, and to handle whatever the small model escalated. Escalation is a first class edge, not an exception. Low confidence comes to me, and my answer becomes input the next time around.

The rebuilt loop: observe, ask the small model, execute or escalate, repeat

The repo I would start from is the open one: Strands Decider on GitHub. Weights, training data, and scripts are all there, so you can fine tune the gates on your own traffic instead of renting someone else’s thresholds.

The sophisticated agent stack this year is not a bigger brain. It is a tiny model welded to every gate, with the big model kept behind it for the work only it can do.



Previous Post
Price by thinking, not by token
Next Post
Three hundred tokens a second