Clef is Cloudflare's open-weights decision model, released on October 1, 2026 and hosted on Workers AI. It takes a state as text, JSON, images or video plus a schema of typed questions and returns a probability for every allowed option in a single forward pass, with no free-form text to parse. It is a 27B model post-trained from Qwen/Qwen3.8-27B with a 64K token context window, licensed under Apache 2.0 and priced at $0.24 per million input tokens on Workers AI. Median request latency is 209.3 ms.
A workable 27B-parameter dense decision model from Cloudflare. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| Acer Veriton GN100 AI MiniAcer | SS | 30.3 GB |
| AMD Instinct MI300XAMD | SS | 30.3 GB |
| AMD Instinct MI325XAMD | SS | 30.3 GB |
| AMD Instinct MI355XAMD | SS | 30.3 GB |
| Apple M3 Ultra (32-core CPU, 80-core GPU)Apple | SS | 30.3 GB |
Cheapest current cloud rentals with at least 30 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 5090Vast.ai · Spot · 32 GB VRAM | $0.20 |
NVIDIA RTX PRO 5000 BlackwellVast.ai · Spot · 48 GB VRAM | $0.25 |
NVIDIA A100 80GB PCIeVast.ai · Spot · 80 GB VRAM | $0.29 |
NVIDIA A40Vast.ai · Spot · 48 GB VRAM | $0.31 |
NVIDIA RTX A6000RunPod · Community · 48 GB VRAM | $0.33 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Clef is Cloudflare's open-weights decision model, designed to eliminate the latency, parsing failures, and non-determinism of standard autoregressive generation in agentic workflows. Built as a dense 27-billion-parameter network post-trained from Qwen/Qwen3.8-27B, Clef bypasses free-form text generation entirely. Instead of outputting tokens sequentially, it reads a state (such as structured JSON, raw text, or multimodal inputs) alongside a typed schema of questions, then returns calibrated probabilities for every permissible option across all questions in a single forward pass.
Released under the Apache 2.0 license, Clef represents a shift toward dedicated "System One" decision components within agent architectures. Where standard large language models require multi-second token generation and fragile JSON schema adherence to make simple triage or routing decisions, Clef produces bounded, typed classification logits directly. Practitioners deploying a local AI model with 27B parameters in 2026 can run Clef locally to handle high-throughput routing, policy enforcement, and classification pipelines without paying the compute tax of autoregressive decoding.
Clef couples a dense 27B transformer backbone with a specialized routing head. The underlying foundation is post-trained from Qwen/Qwen3.8-27B, retaining its core weights and representation capabilities while repurposing the final layers for simultaneous schema evaluation.
Standard LLM classification prompts rely on decoding an answer token by token. Clef replaces the standard language modeling head with a joint schema head:
Because the inference pass stops after evaluating the prompt and computing output logits, Clef avoids the autoregressive decoding phase altogether. This architecture fundamentally decouples decision latency from generation length: whether your schema contains one question or ten, the inference time remains essentially flat.
Clef supports a 65,536-token context window (64K). Dense 27B models usually encounter memory bandwidth bottlenecks when scaled to 64K contexts during generation. However, because Clef only performs prompt ingestion (prefill) and a single forward evaluation step, the memory overhead associated with maintaining a dynamic key-value (KV) cache across thousands of generated tokens is eliminated. The memory footprint scales strictly with the size of the ingested state and the batch size.
Clef is not a conversational chatbot or a code synthesis engine; it is a dedicated decision engine. The model is optimized for high-reliability, low-latency evaluation across discrete tasks:
Because the output is bounded by design, Clef cannot hallucinate keys, break JSON formatting, or produce out-of-distribution values.
To run Clef locally, practitioners must account for the 27B parameter footprint. Because Clef executes only the forward pass without generating downstream tokens, the traditional metric of Clef tokens per second is less relevant than end-to-end forward pass latency. On enterprise hardware like an NVIDIA H200, median latency clocks in at roughly 209 ms. On consumer hardware, performance depends on compute bandwidth during the prefill phase.
Memory requirements depend heavily on quantization. The model checkpoint contains both the backbone weights and the custom joint schema head.
| Quantization Level | Weight Footprint | Minimum VRAM (Context < 4K) | Recommended VRAM (Context up to 64K) | Recommended Target Hardware |
|---|---|---|---|---|
| BF16 / FP16 | ~54 GB | 60 GB | 80 GB+ | 1x NVIDIA A100/H100 (80GB) or 2x RTX 4090 (24GB) via tensor parallelism |
| 8-bit (FP8 / INT8) | ~28 GB | 32 GB | 48 GB | 2x RTX 3090/4090 (24GB) or 1x RTX 6000 Ada (48GB) |
| 4-bit (Q4_K_M / AWQ) | ~16 GB | 19 GB | 24 GB | 1x NVIDIA RTX 3090 / RTX 4090 (24GB) or Apple M-Series (36GB+ Unified Memory) |
For most self-hosted production setups, the best quantization for Clef is FP8 or a high-quality 4-bit format (such as Q4_K_M). Running at 4-bit allows the weights to fit within the 24GB buffer of a single consumer GPU while preserving the calibration of its decision logits.
For engineers wondering how to run 27B model on consumer GPU setups, single-card execution is entirely achievable if you quantize:
The best GPU for Clef in high-density local production environments is a single 48GB card (such as an NVIDIA RTX 6000 Ada or L40S) for 8-bit inference, or an 80GB A100/H100 if running the original BF16 weights at high batch sizes.
Clef requires custom model definition code to handle its joint schema head, meaning standard inference runtimes require specific support:
/v1/systemone. Launching the official SGLang container with --model-path Cloudflare/clef enables batched inference, kernel optimizations, and standard API compatibility.joint_schema_model.py module included in the Hugging Face repository. This requires torch >= 2.11 and transformers >= 5.10.2, loading the model via load_release_model() and collating input records directly in Python.Evaluating Clef hardware requirements and architectural trade-offs requires comparing it directly to its architectural peers and conventional LLMs.
Typesafe AI's Jev pioneered the current System One decision model category. Clef was designed as a drop-in replacement, matching the Jev API request specification (state, questions, and typed schemas for choice, noul, and score).
Practitioners often attempt to achieve structured decisions by prompting general-purpose models using JSON mode or grammar-constrained sampling.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Cloudflare model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.