Full Breakdown
Open-Weight Decision Models Surge as AI Agents Seek Faster, Cost-Effective Choices
By Drooid · · How we work
Core Event: Release of Three Open-Weight Decision Models
On October 1, Cloudflare announced Clef, a pair of open-weight models that return a probability instead of generating text, enabling AI agents to make bounded decisions in milliseconds. The same day, Amazon Web Services released Strands Decider 2B, an open-source 2-billion-parameter model that scores predefined options and reports a confidence score. A week earlier, on September 30, OpenAI introduced its Decisions API, built on the Luna model, to serve the same classification-and-routing use case. All three offerings target the emerging need for “decision” models that replace full-text generation in agent loops.
Background & Context: From Jev to a New Model Category
The race began with TypeSafe AI’s Jev, a model designed for machine-readable decisions. A hackathon demo showed Jev could monitor an AI agent’s actions for $2.94 versus $372 using a frontier LLM, highlighting the cost advantage of decision models. OpenAI and Amazon subsequently entered the space, each building on the Jev concept.
Data & Statistics: Speed, Size, and Benchmark Performance
- Clef: 27 B-parameter “accurate” version and 9 B-parameter “flash” variant; runs on a Qwen backbone with a 64 k-token context window. Cloudflare reports classification of a website in 2.2 s, compared with 4.7 s for gpt-oss-120b, and a median speed 2.5 × faster than Jev across 43 runs. Clef-flash is claimed to be 13 × faster than Jev.
- Strands Decider 2B: Returns a decision in a median of 115 ms on a single Nvidia RTX 3090 and 153 ms on an M3 MacBook. It placed 3rd of 33 models in its size class on JevBench for combined accuracy and calibration, achieving 100 % on JevBench’s easy tasks.
- OpenAI Decisions API: No latency or benchmark figures are provided; the offering is positioned as a hosted service tied to Luna.
Official Statements & Responses
- Cloudflare describes Clef as a “non-autoregressive, prefill-only scoring” approach that skips token-by-token generation and can process up to four images alongside text.
- Amazon calls Strands Decider a “cheap decision” that can replace a full LLM call and notes each decision includes a reliability score not exposed by frontier APIs.
- OpenAI’s Dev Day presentation positioned the Decisions API as the company’s answer to bounded-answer use cases, though no performance metrics were disclosed.
Conflicting Reports & Gaps
- Cloudflare claims Clef outperforms Jev in three of four evaluations on Jev’s benchmark, while Amazon’s results place Strands Decider third in its size class without a direct comparison to Clef.
- No latency or accuracy data are provided for OpenAI’s Decisions API, leaving a gap in understanding how it stacks up against the open-weight alternatives.
- Benchmark methodologies differ across sources, making cross-model performance comparisons uncertain.
Timeline of Key Releases
- September 30: OpenAI launches Decisions API at Dev Day.
- October 1: Cloudflare releases Clef (27 B and 9 B versions).
- October 1, 2026: Amazon Web Services releases Strands Decider 2B.
What’s Next: Future Directions for Decision Models
Both Cloudflare and Amazon indicate plans to integrate their decision models with edge-computing and AI-gateway services, enabling developers to fine-tune and redeploy models via platforms such as Cloudflare Workers AI and AWS’s upcoming libraries for wiring decision models into agents. No specific dates are provided.
