Clef-flash is Cloudflare's 9B multimodal decision model, released on October 1, 2026 and hosted on Workers AI. It reads a state as text, JSON, images or video together with a schema of typed questions and returns a probability for every allowed option in one forward pass, with no free-form text output. It is post-trained from Qwen3.5-9B, open-sourced under Apache 2.0, and priced at $0.038 per million input tokens. Median latency is 38.8 ms.
A workable 9.41B-parameter dense decision model from Cloudflare. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 7.8 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 7.8 GB |
| AMD Instinct MI300XAMD | SS | 7.8 GB |
| AMD Instinct MI325XAMD | SS | 7.8 GB |
| AMD Instinct MI355XAMD | SS | 7.8 GB |
Cheapest current cloud rentals with at least 8 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.06 |
NVIDIA GeForce RTX 5060Vast.ai · Spot · 8 GB VRAM | $0.08 |
NVIDIA RTX A4000Vast.ai · Spot · 16 GB VRAM | $0.08 |
NVIDIA RTX A4000Vast.ai · On-Demand · 16 GB VRAM | $0.09 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Clef-flash is a 9.41B parameter dense decision model built by Cloudflare and released under the Apache 2.0 license. Unlike standard autoregressive large language models that generate free-form text token by token, Clef-flash evaluates an arbitrary input state alongside a schema of typed questions to output calibrated probabilities in a single forward pass. Post-trained from Qwen3.5-9B, the model eliminates the parsing failures, nondeterminism, and high latency typical of prompt-engineered chat models used for classification and routing.
As a local AI model 9.41B parameters 2026 release, Clef-flash establishes an efficient runtime pattern for agentic architectures. Cloudflare designed the model to serve as a fast classifier that determines execution paths, evaluates condition thresholds, and routes context to downstream tools. By computing discrete logits for pre-defined question schemas rather than generating string tokens, Clef-flash delivers median inference latencies around 38.8 ms under dedicated GPU acceleration, making it an attractive engine for high-throughput local pipelines.
Deploying Clef-flash locally gives teams complete data privacy and predictable latency for decision tasks. For engineers building autonomous systems, the model fits between observation ingestion and tool execution, replacing fragile string extraction layers with direct categorical probabilities.
Clef-flash uses a dense transformer architecture totaling 9.41B parameters, built upon the Qwen3.5-9B base checkpoint. Standard LLMs feed output embeddings back through a causal language modeling head sequentially. Clef-flash departs from this mechanism by pairing the frozen or fine-tuned backbone with a dedicated joint schema head.
The joint schema head is a compact transformer module that processes the final hidden states generated by the 9.41B backbone. When a request is submitted, the model encodes both the input state (such as raw text, structured JSON records, or system logs) and a schema of target questions into the context window. The schema head routes contextual representations from the input state directly to the question definitions, calculating logits for every declared option simultaneously.
Key technical characteristics include:
/v1/systemone schema formatBecause the model does not run an autoregressive decode loop, its computational profile mirrors the prefill phase of an LLM. Once the 24,576-token context window is ingested and processed through the dense layers, the head computes output probabilities immediately via a per-question softmax. This prevents KV-cache inflation during token generation and maintains deterministic execution times.
Clef-flash handles classification, scoring, and routing workloads where agents must evaluate structured or unstructured context against strict schemas. Supported schema primitives include boolean checks (noul questions), multi-class categorical selection (choice questions), and continuous numerical evaluations (score questions).
Primary local workloads include:
When you run Clef-flash locally, system resource planning centers entirely on model weight residency and prompt ingestion throughput, rather than autoregressive generation speeds.
Memory requirements depend on the chosen precision and how much of the 24,576-token context window your pipeline populates:
The best GPU for Clef-flash depends on your target concurrency. For standalone developer setups or low-concurrency pipelines, an Nvidia GeForce RTX 4090 (24 GB) is the top consumer option. It runs the full BF16 checkpoint in VRAM with sufficient overhead to handle the complete 24,576 context window without offloading.
For budget-conscious developers asking how to run 9.41B model on consumer GPU hardware with 12 GB or 8 GB VRAM, an RTX 4070 (12 GB) running an 8-bit or 4-bit quantization delivers excellent latency. On Apple Silicon, an M3 Max or M4 Max system with 36 GB or more of unified memory runs the model unquantized with minimal latency overhead.
For the vast majority of local production deployments, Q4_K_M is the best quantization for Clef-flash. Because the model operates as a categorical classifier via its joint schema head, the slight weight degradation of a medium 4-bit k-quantization rarely shifts the argmax decision or significantly warps the calibrated softmax distribution. If your application relies on high-precision probability calibration near tight boundary cutoffs (such as 0.98 vs 0.99), run the model in FP8 or native BF16.
Standard evaluations measure "Clef-flash tokens per second" as a baseline, but that metric can be misleading. Traditional models generate 30 to 90 tokens per second out of the decode phase. Clef-flash does not decode output tokens: it runs one forward pass and halts.
Clef-flash performance is therefore measured in prefill speed and end-to-end request latency:
Ollama and llama.cpp provide direct routes to serve the model locally using the native /v1/systemone format. To run the model inside high-throughput serving stacks, SGLang offers direct backend integration:
1docker run --gpus all \2 --shm-size 32g \3 -p 30000:30000 \4 -v ~/.cache/huggingface:/root/.cache/huggingface \5 --ipc=host \6 lmsysorg/sglang:dev-clef \7 sglang serve \8 --model-path Cloudflare/clef-flash \9 --host 0.0.0.0 \10 --port 30000
Once running, you can dispatch structured JSON payloads directly to the decision endpoint:
1curl http://localhost:30000/v1/systemone \2 -H 'Content-Type: application/json' \3 -d '{4 "model": "Cloudflare/clef-flash",5 "state": "Database connection pool exhausted on worker node 4.",6 "questions": {7 "severity": {8 "type": "choice",9 "instructions": "Determine incident severity level",10 "criteria": {11 "critical": "Service is actively degraded",12 "warning": "Warning threshold reached but functional",13 "info": "Informational state"14 }15 }16 }17 }'
Evaluating Clef-flash vs alternative setups highlights trade-offs in architecture and inference efficiency.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Cloudflare model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.