Cloudflare's Clef-Omni is an open-weight multimodal decision model released on 2026-10-09. It reads a state as text, JSON, images, audio or video and returns a probability for every allowed option of every typed question in a single forward pass, with no free-form text generation or output parsing. It is a 30B-A3B mixture-of-experts model post-trained from Qwen3-Omni-30B-A3B-Instruct, with a 64,000 token context window, and is released under Apache-2.0. Cloudflare hosts it at $0.15 per million input tokens and does not charge for output tokens.
A solid 31.83B-parameter MoE decision model from Cloudflare. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 3.8 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 3.8 GB |
| AMD Instinct MI300XAMD | SS | 3.8 GB |
| AMD Instinct MI325XAMD | SS | 3.8 GB |
| AMD Instinct MI355XAMD | SS | 3.8 GB |
Cheapest current cloud rentals with at least 4 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.06 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA Tesla V100 16GBVast.ai · Spot · 16 GB VRAM | $0.09 |
NVIDIA RTX A4000Vast.ai · Spot · 16 GB VRAM | $0.09 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Clef-Omni is an open-weight 31.83B parameter mixture-of-experts (MoE) decision model released by Cloudflare under the Apache-2.0 license. Built as a specialized post-train of Qwen3-Omni-30B-A3B-Instruct, the architecture routes inference through just 3B active parameters while keeping the full parameter pool available for representation capacity. Unlike standard autoregressive large language models that generate free-form text token by token, Clef-Omni operates as an end-to-end decision engine: it evaluates an input state against a typed schema of questions and produces probability distributions across all predefined choices in a single forward pass.
This operational shift eliminates output parsing, JSON schema validation errors, and the latency overhead of multi-token decoding. By mapping complex input states directly to discrete probability distributions, Clef-Omni addresses high-throughput classification, routing, and scoring tasks that typically require slow, fragile prompt pipelines. It establishes a distinct category in the open-weights ecosystem, functioning less like a conversational agent and more like a high-capacity, multi-task discriminator.
For engineers looking to run Clef-Omni locally, the model presents a unique hardware profile. Because it only computes over 3B parameters per forward pass, inference latency is exceptionally low compared to traditional 30B dense models. However, storing the weights requires allocating enough memory for the full 31.83B parameter footprint. This makes Clef-Omni an ideal candidate for systems where VRAM capacity is sufficient to hold the model weights, allowing users to reap the throughput benefits of sparse execution.
Clef-Omni pairs a modified foundational MoE backbone with a specialized classification head:
The primary architectural advantage is Clef-Omni MoE efficiency. In generative text models, generation speed is bottlenecked by memory bandwidth during auto-regressive decoding. Clef-Omni avoids auto-regressive decoding altogether. A single forward pass through 3B active parameters extracts features and produces decisions immediately, delivering latency numbers that rival small sub-3B dense models while retaining the semantic understanding of a 30B-class network.
Clef-Omni is engineered specifically for deterministic routing, evaluation, and structured classification pipelines:
choice (multi-class selection), score (continuous/graded evaluation), and noul (binary true/false) question types.Because the model outputs normalized probabilities rather than strings, developers can establish programmatic confidence thresholds. If the top probability for a routing decision falls below a calibrated limit (e.g., 0.85), systems can safely escalate the payload to a human operator or a larger reasoning model.
Running a local AI model with 31.83B parameters in 2026 requires understanding the distinction between compute load and memory residency. Clef-Omni requires host or GPU memory large enough to hold all 31.83B parameters, even though computation touches only 3B parameters per forward pass.
Memory requirements depend heavily on your target precision and context utilization:
Because Clef-Omni computes decisions in a single forward pass, traditional metrics like auto-regressive tokens per second do not capture its performance profile. Instead, evaluate the model on forward pass latency:
To run Clef-Omni locally, developers currently rely on the official custom PyTorch implementation provided in the model repository, which includes joint_schema_model.py for handling schema encoding and forward evaluations:
1import sys2import torch3from huggingface_hub import snapshot_download45# Download repository and import custom schema module6path = snapshot_download("Cloudflare/clef-omni")7sys.path.insert(0, path)8from joint_schema_model import load_release_model, encode_record, collate_records910# Load model to GPU (requires ~64GB VRAM for raw bfloat16)11model, processor = load_release_model(path, device="cuda")1213record = {14 "state": "Payment failed due to gateway timeout on node-east.",15 "questions": {16 "severity": {17 "type": "choice",18 "instructions": "Determine operational impact.",19 "criteria": {20 "critical": "Complete outage or transaction loss.",21 "minor": "Non-blocking degradation."22 }23 }24 }25}2627encoded = encode_record(processor.tokenizer, record, processor=processor)28batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))2930with torch.inference_mode():31 logits = model(batch)[0]3233for question, q_logits in zip(encoded.questions, logits):34 probs = q_logits.float().softmax(-1).tolist()35 print(question.question_id, dict(zip(question.option_ids, probs)))
Standard GGUF loaders like pure Ollama or unpatched llama.cpp instances do not support Clef-Omni out of the box due to its custom joint schema transformer head. Practitioners using Ollama should look for community backports or run the lightweight PyTorch/Transformers server containerized alongside their local services.
Dense generative models like Qwen2.5-32B or traditional MoEs like Mistral 8x7B generate responses sequentially. When tasked with structured extraction, generative models require prompt constraints, grammars (like Outlines or GBNF), and multi-step sampling that consumes compute for every generated token. Clef-Omni trades open-ended text generation for absolute schema adherence and sub-100ms latency. If your pipeline needs creative summaries or free-form reasoning text, choose a standard LLM. If your pipeline needs to route, classify, or score data according to fixed criteria, Clef-Omni is significantly faster and computationally cheaper.
Traditional cross-encoders like DeBERTa-v3-Large are compact (under 500M parameters) and run on entry-level GPUs. However, they lack the deep semantic comprehension, complex instruction-following, and massive 64,000-token context windows found in modern 30B foundation models. Clef-Omni provides the contextual capacity and complex reasoning of a frontier base model while preserving the single-pass classification dynamics of an encoder head. The tradeoff is memory: Clef-Omni requires at least 20 GB of VRAM, whereas small cross-encoders run comfortably inside 2 GB to 4 GB.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Cloudflare model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.