Drooid Logo
Back to story perspectives

Full Breakdown

Open-Weight Decision Models Surge as AI Agents Seek Faster, Cost-Effective Choices

By Drooid · · How we work

Core Event: Release of Three Open-Weight Decision Models

On October 1, Cloudflare announced Clef, a pair of open-weight models that return a probability instead of generating text, enabling AI agents to make bounded decisions in milliseconds. The same day, Amazon Web Services released Strands Decider 2B, an open-source 2-billion-parameter model that scores predefined options and reports a confidence score. A week earlier, on September 30, OpenAI introduced its Decisions API, built on the Luna model, to serve the same classification-and-routing use case. All three offerings target the emerging need for “decision” models that replace full-text generation in agent loops.

Background & Context: From Jev to a New Model Category

The race began with TypeSafe AI’s Jev, a model designed for machine-readable decisions. A hackathon demo showed Jev could monitor an AI agent’s actions for $2.94 versus $372 using a frontier LLM, highlighting the cost advantage of decision models. OpenAI and Amazon subsequently entered the space, each building on the Jev concept.

Data & Statistics: Speed, Size, and Benchmark Performance

  • Clef: 27 B-parameter “accurate” version and 9 B-parameter “flash” variant; runs on a Qwen backbone with a 64 k-token context window. Cloudflare reports classification of a website in 2.2 s, compared with 4.7 s for gpt-oss-120b, and a median speed 2.5 × faster than Jev across 43 runs. Clef-flash is claimed to be 13 × faster than Jev.
  • Strands Decider 2B: Returns a decision in a median of 115 ms on a single Nvidia RTX 3090 and 153 ms on an M3 MacBook. It placed 3rd of 33 models in its size class on JevBench for combined accuracy and calibration, achieving 100 % on JevBench’s easy tasks.
  • OpenAI Decisions API: No latency or benchmark figures are provided; the offering is positioned as a hosted service tied to Luna.

Official Statements & Responses

  • Cloudflare describes Clef as a “non-autoregressive, prefill-only scoring” approach that skips token-by-token generation and can process up to four images alongside text.
  • Amazon calls Strands Decider a “cheap decision” that can replace a full LLM call and notes each decision includes a reliability score not exposed by frontier APIs.
  • OpenAI’s Dev Day presentation positioned the Decisions API as the company’s answer to bounded-answer use cases, though no performance metrics were disclosed.

Conflicting Reports & Gaps

  • Cloudflare claims Clef outperforms Jev in three of four evaluations on Jev’s benchmark, while Amazon’s results place Strands Decider third in its size class without a direct comparison to Clef.
  • No latency or accuracy data are provided for OpenAI’s Decisions API, leaving a gap in understanding how it stacks up against the open-weight alternatives.
  • Benchmark methodologies differ across sources, making cross-model performance comparisons uncertain.

Timeline of Key Releases

  • September 30: OpenAI launches Decisions API at Dev Day.
  • October 1: Cloudflare releases Clef (27 B and 9 B versions).
  • October 1, 2026: Amazon Web Services releases Strands Decider 2B.

What’s Next: Future Directions for Decision Models

Both Cloudflare and Amazon indicate plans to integrate their decision models with edge-computing and AI-gateway services, enabling developers to fine-tune and redeploy models via platforms such as Cloudflare Workers AI and AWS’s upcoming libraries for wiring decision models into agents. No specific dates are provided.