Holo4-35B-A3B is an agentic computer-use model from Hcompany, released September 28, 2026. It is a Mixture of Experts model with 35B total parameters and 3B active parameters. It works through GUIs, code, MCP and APIs, running on desktop, web, Android and code sandboxes, and is available on the H Models API. It scores 30.9% on OSWorld 2.0.
A solid 35B-parameter MoE language model from Hcompany. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 1.7 GB | Low | |
| Q4_K_MRecommended | 2.3 GB | Good | |
| Q5_K_M | 2.6 GB | Very Good | |
| Q6_K | 3.0 GB | Excellent | |
| Q8_0 | 3.7 GB | Near Perfect | |
| FP16 | 6.6 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| AMD Radeon RX 7600 8GBAMD | SS | 99.1 tok/s | 2.3 GB |
| NVIDIA GeForce RTX 4060NVIDIA | SS | 93.6 tok/s | 2.3 GB |
| NVIDIA GeForce RTX 5060 Ti 8GBNVIDIA | SS | 154.2 tok/s | 2.3 GB |
| AMD Radeon RX 7700 XTAMD | SS | 148.7 tok/s | 2.3 GB |
| Intel Arc B580Intel | SS | 157.0 tok/s | 2.3 GB |
Energy cost on Raspberry Pi 5 (8GB) (~12 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Holo4-35B-A3B on Raspberry Pi 5 (8GB) · ~12 tok/s · 12W | $0.034 |
GPT-6 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 2 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.04 |
NVIDIA GeForce RTX 2080 TiVast.ai · Spot · 11 GB VRAM | $0.04 |
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.04 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 5060Vast.ai · Spot · 8 GB VRAM | $0.05 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cumulative cost of buying the smallest device in our directory that runs this model at 4-bit, compared with paying flagship APIs for the same token volume.
Assumes 1 million tokens a day, a 70/30 input and output mix and $0.12 per kWh. The device only runs as long as the workload needs. Use the ROI calculator for your own numbers.
Holo4-35B-A3B is Hcompany's second model in the Holo4 series, released September 28, 2026, and it is built for one job: driving software the way a person does. It clicks and types on a screen, writes and executes its own code, and calls MCP or API tools, choosing whichever interface fits the task in front of it. That combination is rarer than it sounds. Most agentic models commit to a single interface. GUI-first models are useless when a task has a clean API, and tool-calling specialists stall in front of a desktop application that never exposed one. Holo4 was trained across desktop, web, Android, and code sandboxes specifically so a single workflow can mix all of them.
The architecture is a Mixture of Experts design with 35B total parameters and 3B active per token. That ratio is the headline for anyone planning to self-host. You pay the memory cost of a 35B model but the compute cost of roughly a 3B one, which means generation speed on local hardware is far better than the parameter count suggests. Hcompany ships it alongside a 27B dense sibling and Holotron4 Nano, and publishes weights in FP16, FP8, and GGUF formats, so this is a model you can genuinely run on your own machine rather than only through the H Models API.
Where it stands out is computer use. It scores 30.9% on OSWorld 2.0, a harder successor to the original OSWorld suite that Hcompany used to report 85.2% for the dense 27B model, so treat those two numbers as separate benchmarks rather than a ranking. If your workload involves operating real applications, navigating GUIs, or orchestrating mixed tool and screen interactions, this is the category to evaluate. If you just need a chat model, you are paying for capabilities you will not use.
Holo4-35B-A3B is a sparse Mixture of Experts transformer. Of the 35B total parameters, only about 3B are activated for any given token, routed through a subset of expert layers. The practical consequences split in two directions:
This is the single most common mistake people make when sizing hardware for MoE models. A 3B active parameter count does not mean 3B of memory. Budget for the full 35B weight footprint plus KV cache.
Context length is not specified in the published model card, and neither is the license or training cutoff. Check the Hugging Face repository before committing to a deployment, particularly if you need a permissive license for commercial use or a defined context window for long-document agent traces.
The model is multimodal: text and image input, with vision feeding directly into the computer-use loop. Screen understanding is not a bolt-on here, it is the primary input path.
The capability set is code, reasoning, function-calling, instruction-following, and vision. In practice that maps to a specific kind of work:
It is not positioned as a long-form writing model or a general chat assistant. Evaluate it against agentic workloads.
Because GGUF weights are published, local inference is a realistic path. The constraints are memory bandwidth and capacity, not raw compute.
| Quantization | Approx. weight size | Notes |
|---|---|---|
| FP16 | ~70 GB | Two 48GB cards or a 96GB+ workstation GPU |
| FP8 | ~35 GB | Single 48GB card, or 2x24GB with offload |
| Q8_0 | ~37 GB | Near-lossless, needs 48GB |
| Q6_K | ~29 GB | Good quality, 32GB card or 48GB unified memory |
| Q5_K_M | ~25 GB | Sweet spot for 32GB cards |
| Q4_K_M | ~21 GB | Recommended for most users |
| Q3_K_M | ~18 GB | Fits 24GB with room for context, noticeable quality drop |
Add roughly 2GB-8GB for KV cache depending on context length and batch size.
If you are asking how to run a 35B model on a consumer GPU, this is one of the better candidates in 2026 precisely because the active parameter count keeps generation fast.
These are ballpark figures for Q4_K_M with a moderate context, and they depend heavily on backend and batch size:
Prompt processing on vision input is the slower phase. Budget for it if your agent takes frequent screenshots.
Ollama is the fastest route to a working local instance once the GGUF quants are pulled from the Holo4 collection. Install Ollama, pull the model tag, and run it. For anything beyond a quick test, drop to llama.cpp or vLLM so you can control context size, KV cache quantization, and expert offload explicitly. If you want the hosted equivalent for comparison, Hcompany exposes an OpenAI-compatible endpoint at https://api.hcompany.ai/v1/ with the model name holo4-35b-a3b, so an existing OpenAI client works with a two-line change.
Against Holo4-27B (dense, same family). The 27B dense model activates all 27B parameters per token, so it is slower and needs a comparable memory budget, but dense models often hold up better on tasks requiring tight, consistent reasoning across long chains. Choose the 27B if your workload is mostly code and text reasoning with occasional screen interaction. Choose the 35B-A3B if throughput matters and computer use is central.
Against Qwen3-30B-A3B. The closest architectural peer: a 30B MoE with roughly 3B active parameters. Qwen3 is stronger as a general-purpose text and coding model and has a well-documented context window and license. Holo4-35B-A3B is the better pick when the task involves operating a GUI or mixing screen, code, and tool interfaces in one loop. For pure text generation, Qwen3 is usually the more predictable choice.
The deciding question is whether you need an agent that can see and drive a screen. If yes, Holo4-35B-A3B is one of the few open-weight options in this size class. If no, a dense model or a general MoE will serve you better for the same memory budget.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Hcompany model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.