Hcompany released Holo4-27B on 2026-09-28, a dense 27B agentic model built for computer use. It works through GUIs, code, MCP and API tools, and runs on desktop, web, Android and code sandboxes. Weights are published in FP16, FP8 and GGUF formats, and the model is also served on the H Models API. It scores 61.7% on OSWorld 2.0.
A solid 27B-parameter dense language model from Hcompany. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 11.4 GB | Low | |
| Q4_K_MRecommended | 17.0 GB | Good | |
| Q5_K_M | 19.7 GB | Very Good | |
| Q6_K | 23.0 GB | Excellent | |
| Q8_0 | 29.7 GB | Near Perfect | |
| FP16 | 55.4 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| AMD Radeon RX 7900 XTXAMD | SS | 45.3 tok/s | 17.0 GB |
| Google Cloud TPU v6e (Trillium)Google | SS | 77.5 tok/s | 17.0 GB |
| NVIDIA GeForce RTX 3090NVIDIA | SS | 44.2 tok/s | 17.0 GB |
| NVIDIA GeForce RTX 4090 Founders EditionNVIDIA | SS | 47.6 tok/s | 17.0 GB |
| NVIDIA GeForce RTX 5090 Founders EditionNVIDIA | SS | 84.6 tok/s | 17.0 GB |
Energy cost on Intel Arc A770 16GB (~26 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Holo4-27B on Intel Arc A770 16GB · ~26 tok/s · 225W | $0.284 |
GPT-6 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 17 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3090Vast.ai · Spot · 24 GB VRAM | $0.09 |
NVIDIA RTX PRO 4000 BlackwellVast.ai · Spot · 24 GB VRAM | $0.13 |
NVIDIA RTX A6000Vast.ai · Spot · 48 GB VRAM | $0.13 |
NVIDIA L4Vast.ai · Spot · 24 GB VRAM | $0.13 |
NVIDIA GeForce RTX 3090Vast.ai · On-Demand · 24 GB VRAM | $0.14 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cumulative cost of buying the smallest device in our directory that runs this model at 4-bit, compared with paying flagship APIs for the same token volume.
Assumes 1 million tokens a day, a 70/30 input and output mix and $0.12 per kWh. The device only runs as long as the workload needs. Use the ROI calculator for your own numbers.
Holo4-27B is Hcompany's dense 27B agentic model built for computer use. Released on 2026-09-28, it is designed to drive software the way a person does: clicking and typing through GUIs, writing and executing code, and calling MCP and API tools. It runs across desktop, web, Android, and code sandboxes, and Hcompany publishes weights in FP16, FP8, and GGUF formats alongside an OpenAI-compatible endpoint on the H Models API.
The distinguishing design choice is interface flexibility. Most agentic models specialize: GUI-trained models are useless when there is no screen, and tool-calling models stall in front of an application with no API. Holo4-27B was trained to pick whichever interface the task demands, which matters for real workflows where a single job spans a browser, a legacy desktop app, and a shell. It scores 61.7% on OSWorld 2.0, the benchmark Hcompany uses to measure generalist computer-use performance.
For local deployment, the dense 27B form factor is the pragmatic part. Every parameter is active on every token, so VRAM planning is straightforward and there is no expert-routing behavior to profile. It sits in the same weight class as Qwen3-VL-32B and Gemma 3 27B, but its training target is agentic control rather than general chat or document understanding.
Holo4-27B is a dense transformer with 27B parameters, paired with a vision encoder for screenshot and image input. Because it is dense rather than Mixture-of-Experts, all 27B parameters are read on every forward pass. That has two practical consequences:
Hcompany has not published a context length for Holo4-27B, and no training cutoff or license terms were specified at release. Treat context as an open question: verify the current config on the model card before you architect a long-horizon agent around it, and budget KV cache accordingly. For computer-use loops, where each step carries a screenshot plus an action history, KV cache grows faster than in plain chat, so headroom above the weight footprint matters.
The sibling model, Holo4-35B-A3B, is a Mixture-of-Experts variant with 3B active parameters. It is faster per token but needs the full 35B resident in memory. Holo4-27B trades that speed for a smaller total footprint and more consistent latency.
The model covers chat, code, vision, and function calling, and the useful work happens where those overlap.
If you want a general-purpose chat model or a long-context document reasoner, this is not the sharpest tool. If you want an agent that operates software, it is built for exactly that.
GGUF weights are published, which means the standard local stack applies: llama.cpp, Ollama, and LM Studio for GGUF builds, and vLLM or SGLang for FP8 and FP16 on server-class GPUs.
| Precision | Approx. weights | Realistic minimum VRAM |
|---|---|---|
| Q3_K_M | ~13.5 GB | 16 GB |
| Q4_K_M | ~16.5 GB | 20-24 GB |
| Q5_K_M | ~19 GB | 24 GB |
| Q6_K | ~22 GB | 24-32 GB |
| Q8_0 | ~28.7 GB | 32-40 GB |
| FP8 | ~27 GB | 32-48 GB |
| FP16 | ~54 GB | 64-80 GB |
Add 2-6 GB on top for KV cache and the vision encoder, more if you run long agent trajectories or high concurrency.
Q4_K_M is the right default. It fits a single 24 GB card with room for context, and the quality loss against FP16 is small enough that agentic action accuracy barely moves. Step up to Q5_K_M or Q6_K only if you have a 32 GB card and want the extra margin on vision-heavy tasks. FP8 is the better choice than Q8_0 if your runtime supports it, since it uses less memory for comparable quality. FP16 is for 80 GB accelerators or multi-GPU setups and is not worth the bandwidth cost on consumer hardware.
For a 27B dense model on a consumer GPU, 24 GB is the practical floor and 32 GB is the comfortable target.
Ollama is the fastest way to get started, assuming the GGUF tag is published in the Holo4 collection:
1ollama run holo4:27b
Then point your agent loop at the local endpoint. For vision and function calling, confirm your runtime supports image input and tool-call parsing for this architecture before you build on it, since support varies between llama.cpp builds.
Holo4-27B vs Qwen3-VL-32B. Qwen3-VL-32B is the closest general-purpose alternative: dense, vision-capable, slightly larger, and stronger on document understanding and broad multimodal chat. Choose Qwen3-VL-32B if your workload is mostly reading and reasoning over images. Choose Holo4-27B if your workload is acting on screens, where its training on GUI trajectories and tool interfaces gives it a clear edge.
Holo4-27B vs Gemma 3 27B. Gemma 3 27B is a solid general chat and vision model with a permissive license and broad runtime support, and it is easier to deploy. It is not an agent model. If you need something to drive a browser or fill a form, Holo4-27B is the correct pick; if you need a conversational assistant with image input, Gemma 3 27B is lighter to operate and better documented.
Holo4-27B vs Holo4-35B-A3B. Same family, different tradeoff. The MoE model is faster per token but requires 35B of resident memory. On a single 24 GB card, the dense 27B at Q4_K_M is the only one of the two that fits.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Hcompany model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.