Every open model in our directory, ranked by how well it runs on HP ZGX Fury G1n AI Station (252GB) at 4-bit. Grades weigh whether the model fits in memory and how fast it should run.
HP ZGX Fury G1n AI Station has 252 GB of VRAM, enough to run 275 of the 288 open models we track at 4-bit. The largest that still fits well is Kimi K2.7 Code (1000B), and the top-rated is Naive-N0.5-Flash at about 43 tok/s.
275 of 288 models run comfortably on HP ZGX Fury G1n AI Station.
See full HP ZGX Fury G1n AI Station specs and pricingThe top models this device can run at 4-bit, ranked by fit and speed.
| Model | Grade | Speed | VRAM |
|---|---|---|---|
| Naive-N0.5-FlashNaiveAI | SS | 42.8 tok/s | 133.5 GB |
| MiMo-V2.6-Flash-RLXiaomi MiMo | SS | 44.2 tok/s | 129.2 GB |
| DeepSeek-V4.1-Flashdeepseek-ai | SS | 41.5 tok/s | 137.8 GB |
| Llama 4 MaverickMeta | SS | 39.1 tok/s | 146.4 GB |
| Llama 3.1 70B InstructMeta | SS | 50.7 tok/s | 112.8 GB |
| Llama 3.3 70B InstructMeta | SS | 50.7 tok/s | 112.8 GB |
| DeepSeek-V4-FlashDeepSeek | SS | 51.0 tok/s | 112.0 GB |
| Nvidia Nemotron 3 SuperNVIDIA | SS | 55.2 tok/s | 103.5 GB |
| GLM-5Z.ai | SS | 65.2 tok/s | 87.7 GB |
| GLM-5.1Z.ai | SS | 65.2 tok/s | 87.7 GB |
| Kimi K2.7 CodeMoonshot AI | SS | 66.3 tok/s | 86.2 GB |
| Kimi K2.6Moonshot AI | SS | 66.3 tok/s | 86.2 GB |
Naive-N0.5-FlashNaiveAI | 309B(15.5B active) | SS | 42.8 tok/s | 133.5 GB | |
MiMo-V2.6-Flash-RLXiaomi MiMo | 310.76B(15B active) | SS | 44.2 tok/s | 129.2 GB | |
DeepSeek-V4.1-Flashdeepseek-ai | 552B(16B active) | SS | 41.5 tok/s | 137.8 GB | |
Llama 4 MaverickMeta | 400B(17B active) | SS | 39.1 tok/s | 146.4 GB | |
| 70B | SS | 50.7 tok/s | 112.8 GB | ||
| 70B | SS | 50.7 tok/s | 112.8 GB | ||
DeepSeek-V4-FlashDeepSeek | 284B(13B active) | SS | 51.0 tok/s | 112.0 GB | |
Nvidia Nemotron 3 SuperNVIDIA | 120B(12B active) | SS | 55.2 tok/s | 103.5 GB | |
| Ad | |||||
GLM-5Z.ai | 744B(40B active) | SS | 65.2 tok/s | 87.7 GB | |
GLM-5.1Z.ai | 744B(40B active) | SS | 65.2 tok/s | 87.7 GB | |
Kimi K2.7 CodeMoonshot AI | 1000B(32B active) | SS | 66.3 tok/s | 86.2 GB | |
Kimi K2.6Moonshot AI | 1000B(32B active) | SS | 66.3 tok/s | 86.2 GB | |
Kimi K2 Instruct 0905Moonshot AI | 1000B(32B active) | SS | 67.6 tok/s | 84.6 GB | |
Kimi K2 ThinkingMoonshot AI | 1000B(32B active) | SS | 67.6 tok/s | 84.6 GB | |
Kimi K2.5Moonshot AI | 1000B(32B active) | SS | 67.6 tok/s | 84.6 GB | |
Falcon 180BTechnology Innovation Institute | 180B | SS | 53.0 tok/s | 107.8 GB | |
| Ad | |||||
IQuest-Q1IQuest | 320.32B(15B active) | SS | 79.3 tok/s | 72.1 GB | |
GLM-4.6Z.ai | 355B(32B active) | SS | 81.3 tok/s | 70.3 GB | |
Ring-2.6-1TinclusionAI | 1000B(63B active) | SS | 33.8 tok/s | 169.2 GB | |
Gemma 4 31B ITGoogle | 31B | SS | 69.7 tok/s | 82.0 GB | |
Mistral Large 3 675BMistral AI | 675B(41B active) | SS | 86.3 tok/s | 66.3 GB | |
DeepSeek-V3DeepSeek | 671B(37B active) | SS | 95.5 tok/s | 59.8 GB | |
DeepSeek-R1DeepSeek | 671B(37B active) | SS | 95.5 tok/s | 59.8 GB | |
DeepSeek-V3.1DeepSeek | 671B(37B active) | SS | 95.5 tok/s | 59.8 GB | |
| Ad | |||||
DeepSeek-V3.2DeepSeek | 685B(37B active) | SS | 95.5 tok/s | 59.8 GB | |

Model Compatibility
Use the calculator to weigh this device against any other, model by model, with speed and fit scores.
Straight answers to the questions we hear most often.
Can't find what you're looking for? Book a Discovery Call
Naive-N0.5-Flash is the highest-rated model that runs on HP ZGX Fury G1n AI Station at 4-bit, at about 43 tok/s. In total, HP ZGX Fury G1n AI Station can run 275 of the 288 open models we track.
We size each model at 4-bit (Q4_K_M), add the KV cache for its context length and a fixed runtime overhead, then compare that to this device VRAM. Each model is graded on how comfortably it fits and how fast it should run on this card.
It is an estimate of decode speed in tokens per second at 4-bit, based on this device memory bandwidth and each model size. Real speed depends on your runtime, batch size, and context length, so use it to compare models, not as a hard promise.
Lower-bit quantization shrinks a model so more of them fit. These grades already assume 4-bit, which is the common local default. Going below 4-bit can squeeze a larger model in at some quality cost; going above needs more VRAM and may not fit.
Mixture-of-experts models only activate a fraction of their parameters at a time, so their runtime memory is closer to the active size than the headline size. We score fit on the active parameters, which is why some very large models still fit.