Cloudflare Clef vs Strands Decider: Which Decision Model Should Your AI Agent Use?
By Faisal Khan

On October 1, Cloudflare and AWS shipped the same new kind of model on the same day. Cloudflare released Clef. AWS's Strands Labs released Strands Decider 2B. Both are "decision models": they don't write text, they pick from options you give them and tell you how sure they are.
If you build AI agents, this matters more than another chat model launch.
My short answer on Cloudflare Clef vs Strands Decider: pick Clef if you want a hosted API with vision and the stronger published scores. Pick Strands Decider if you want something small, free, and running on your own machine. Either one can take routine decisions away from an expensive LLM.
What are Cloudflare Clef and Strands Decider?
Cloudflare Clef and Strands Decider are decision models released on October 1, 2026: small models that read a situation plus a list of typed questions and return a probability for each allowed answer instead of generating text. Cloudflare Clef is hosted on Workers AI, while AWS's Strands Decider 2B is built to run locally.
The category is new. Both companies credit TypeSafe AI's Jev, released in September, for kicking it off. TechCrunch's report on the Strands launch describes dozens of Jev-style models appearing within weeks, and the Strands release itself as Amazon's own take on the idea.
A decision model answers questions like these:
- Is this support message urgent? Yes or no, with a probability.
- Which team should handle it: billing, technical, or sales?
- Is the agent about to call a tool with an argument the user never gave?
That's the whole job. No paragraphs, no summaries. Just a typed answer your code can act on.
How are Clef and Strands Decider built differently?
Clef and Strands Decider both start from a Qwen language model and strip out its ability to write text, but they differ a lot in size and where they run. Clef is a 27-billion-parameter model on Cloudflare's network. Strands Decider is a 2-billion-parameter model meant for your laptop or server.
According to Cloudflare's launch post, Clef freezes Qwen3.8-27B (Qwen3.5-9B for Clef-flash) and trains a routing head plus low-rank adapters on top. Each request does one pass over the input, then scores every valid answer in parallel. Clef also has a vision encoder, so it can classify images, and a 65,536-token context window.
AWS's Strands post takes Qwen3.5-2B, removes the language-model head, and adds a pointer head of just over a million parameters that scores each option. AWS published the training data and scripts in the strands-decider GitHub repo, along with the weights. Honestly, that openness is the most interesting part. You can read every version they tried, up to v19.
Clef's weights are also open, on Hugging Face under Apache 2.0. The difference is that running a 27B model yourself needs a serious GPU. Most people will use it through the API.
Cloudflare Clef vs Strands Decider: specs and pricing side by side
Clef costs money per token but needs no hardware, while Strands Decider is free to use but runs on hardware you pay for. The table below pulls the published numbers from each vendor's own pages.
| Cloudflare Clef | Clef-flash | Strands Decider 2B | |
|---|---|---|---|
| Size | 27B | 9B | 2B |
| Where it runs | Workers AI API (or self-host) | Workers AI API (or self-host) | Your CPU, GPU or Mac |
| Price | $0.24 per million input tokens | $0.09 per million input tokens | Free, you pay for hardware |
| Median latency | 209.3ms | 38.8ms | about 115ms (RTX 3090), about 153ms (M3 MacBook) |
| Images | Yes, up to 4 per request | Yes, up to 4 per request | No |
| Licence | Apache 2.0 | Apache 2.0 | Apache 2.0 |
| Best fit | Accuracy-sensitive hosted decisions | High-volume, latency-critical decisions | Local dev, privacy, zero per-call cost |
Prices come from Cloudflare's model pages for Clef and Clef-flash. Clef's latency figures are Cloudflare's own medians across its benchmark runs. Strands' figures are AWS's, measured on small tasks, and AWS notes latency grows roughly in line with input size.
Here's a quick cost check, using my own arithmetic from list prices. Say each decision sends about 500 input tokens, and your agent makes 10,000 decisions a day. That's 5 million tokens a day: about $1.20 on Clef, about $0.45 on Clef-flash. Over a month, roughly $36 and $13.50.
Which one scores better?
Clef has the stronger published results, but the two vendors measured on different tests, so there is no clean head-to-head yet. Treat every number here as a vendor claim until independent tests catch up.
Cloudflare compared Clef against Jev and other decision models on ten benchmarks. Some highlights from its table: Clef scored 94.20 macro-F1 on BANKING77 against Jev's 79.74, and 98.47 on BFCL case-exact against Jev's 95.75. Jev still won on When2Call (80.97 against Clef's 72.37) and on agent trace observability. Strands Decider isn't in Cloudflare's table.
AWS took a different route. It reports Strands Decider as 3rd of 33 models in the 2B class on JevBench's public set for accuracy and calibration, and says it gets 100% of JevBench's easy tasks right. AWS is also upfront about the limits: decision models are "significantly worse at solving complex problems than reasoning models."
That honesty is useful. A 2B model will miss harder calls that a 27B model gets right. The question is whether your decisions are hard.
Why should an AI agent use a decision model at all?
An AI agent should use a decision model for the small, repeated yes-or-no and pick-one choices it makes on every run, because a decision model answers those faster, cheaper, and with a confidence score a general LLM API doesn't give you.
Think about what a typical agent does between real tasks. Which tool fits this request? Is this message spam? Does this action need human approval? Each of those is usually a full LLM call today. That adds cost and a second or two of waiting, every single time.
Cloudflare gives one real example. Its threat intelligence team used Clef to classify website domains: 2.2 seconds to fetch, render, and classify, against 4.7 seconds for gpt-oss-120b in the same workflow.
The confidence score matters as much as the speed. The Strands repo includes an example where the decider checks a proposed tool call before it runs: are the arguments grounded in what the user actually said, and is it too early to call the tool? If confidence is low, the agent asks the user instead of guessing. That's exactly the kind of guardrail I cover in AI agent security, and now it costs milliseconds instead of a full model call.
How to choose between Clef and Strands Decider
Choose by where your agent runs and how hard your decisions are, then confirm with your own data. These are the steps I'd follow.
- List the decisions your agent makes. Routing, tool choice, approval checks, spam filters. Write each one as a question with fixed answers.
- Check if images are involved. If a decision needs to look at a screenshot, photo, or scanned document, only Clef handles that today.
- Decide where the model can live. Already on Cloudflare Workers? Clef is one API call away. Need data to stay on your own servers, or zero per-call cost? Start with Strands Decider.
- Pull 50–100 real examples per decision. Take them from your logs with personal data removed, and label the right answer.
- Run each model and compare accuracy and confidence. Look at the wrong answers. A model that is wrong but unsure is far safer than one that is wrong and confident.
- Keep an LLM fallback. Send low-confidence decisions to your main LLM or a human. The decision model handles the easy majority, and the hard cases still get proper reasoning.
Fitting these models into a real agent, with thresholds and fallbacks, is the kind of work I do in AI agent development. The model itself is the easy part. Deciding what happens when it's unsure is where agents succeed or fail.
Is Cloudflare Clef better than Strands Decider?
Cloudflare Clef has stronger published benchmark scores and supports images, so it is the better choice for accuracy-sensitive hosted decisions. Strands Decider is better when you need a small model that runs locally for free. The vendors used different tests, so neither result is a direct head-to-head comparison.
Can a decision model replace an LLM in my agent?
A decision model cannot replace an LLM in an agent, because it cannot generate text, write code, or reason through complex problems. It can replace the many small LLM calls an agent makes to route requests, pick tools, and check actions, leaving the main LLM for the work that needs real reasoning.
How much does Cloudflare Clef cost?
Cloudflare Clef costs $0.24 per million input tokens on Workers AI, and Clef-flash costs $0.09 per million input tokens, with no charge for output tokens. A decision of about 500 input tokens costs a small fraction of a cent, so even ten thousand decisions a day stays around a dollar.
Can I run Strands Decider 2B on a laptop?
Strands Decider 2B can run on a laptop, since it has about two billion parameters and runs on a CPU, a consumer GPU, or an Apple silicon Mac. AWS reports a median of about 153 milliseconds per small decision on an M3 MacBook, and about 115 milliseconds on an RTX 3090.
Where to start
Cloudflare Clef vs Strands Decider comes down to hosting: Clef for a fast, accurate API with vision, Strands Decider for a free local model you fully control. The bigger win is using either one at all, so your agent stops paying LLM prices for yes-or-no questions. If you want help fitting a decision model into your agent, get in touch.
