MF-1 (Multimodal Flow) is a unified language and vision model from Horizon Robotics with Huazhong University of Science and Technology and Beijing Jiaotong University, released on 2026-10-04. It models text and images in one shared continuous representation with Flow Matching instead of separate encoders or discrete image tokens, and covers language modeling, visual understanding, image generation and editing. The largest version has 1.6B parameters and was trained on 150 billion pre-training tokens. It is fully open and scores 82.8 on GenEval and 75.3 on DPG-Bench.
A solid 1.6B-parameter dense image generator from Horizon Robotics. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 1.5 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 1.5 GB |
| AMD Instinct MI300XAMD | SS | 1.5 GB |
| AMD Instinct MI325XAMD | SS | 1.5 GB |
| AMD Instinct MI355XAMD | SS | 1.5 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 4060 TiVast.ai · Spot · 8 GB VRAM | $0.06 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA Tesla V100 16GBVast.ai · On-Demand · 16 GB VRAM | $0.08 |
NVIDIA RTX A4000Vast.ai · Spot · 16 GB VRAM | $0.08 |
NVIDIA RTX A4000Vast.ai · On-Demand · 16 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
MF-1 (Multimodal Flow) is an open-weights, 1.6-billion-parameter foundation model developed by Horizon Robotics in collaboration with Huazhong University of Science and Technology and Beijing Jiaotong University. Released in October 2026, the model departs fundamentally from the standard design pattern of modern multimodal systems. Rather than bolting a discrete vision transformer to an autoregressive text model, MF-1 establishes a unified representation for both language and images using continuous Flow Matching.
The model is built to handle language modeling, visual question answering, text-to-image generation, and image editing natively within a single set of weights. Trained on 150 billion pre-training tokens, MF-1 punches far above its parameter class, scoring 82.8 on GenEval and 75.3 on DPG-Bench. For developers seeking a capable local AI model with 1.6B parameters in 2026, MF-1 represents an efficient, all-in-one vision-language engine capable of both multimodal comprehension and generation.
Because it eliminates separate vision encoders and discrete vector-quantization bottlenecks, MF-1 presents unique performance characteristics for on-device deployment. Its compact footprint makes it an ideal candidate for edge devices, consumer workstations, and embedded robotics hardware where memory bandwidth and capacity are constrained.
MF-1 uses a dense 1.6-billion-parameter architecture built on a chunk-causal flow backbone. Most vision-language models rely on a dual-component pipeline: a visual encoder (such as CLIP or SigLIP) compresses pixels into embeddings, which are then projected into a causal transformer backbone alongside text tokens. Systems that generate images typically reverse this via discrete visual tokens (VQ-VAE/VQ-GAN) or pipe the output into a separate diffusion model.
MF-1 replaces this disjointed setup with a continuous generative flow:
Because the model is dense, every parameter is engaged during a forward pass. This gives MF-1 consistent computational characteristics without the routing overhead or uneven memory spikes associated with sparse Mixture-of-Experts (MoE) architectures.
MF-1 handles both perception and generation tasks within one weight distribution. The training recipe enables several key capabilities:
Due to its 1.6B dense footprint, the MF-1 hardware requirements are remarkably modest, making it accessible on almost any modern consumer GPU, integrated SoC, or older enterprise accelerator.
The precise memory footprint depends on the execution mode (pure language generation vs. multi-step flow integration for images) and the quantization format:
To run MF-1 locally, practically any machine with a dedicated GPU or unified memory will suffice:
Choosing the best quantization for MF-1 depends on whether your workload is text-centric or vision-centric:
When deployed on an NVIDIA RTX 4070 (12GB), expect the following MF-1 performance figures:
If you are looking at how to run 1.6B model on consumer GPU setups via standard frameworks, containerized runtimes and execution backends supporting flow matching layers provide the cleanest path. For standard textual inference, Ollama and llama.cpp offer straightforward execution once model weights are converted to GGUF format, though full multimodal image synthesis requires runtimes that implement the model's unified flow matching ODE/SDE solvers.
Evaluating MF-1 against alternative compact architectures highlights the trade-offs of its unified continuous flow design:
Qwen2.5-VL-1.5B is a standard-bearer in the sub-2B vision-language tier. It pairs a discrete ViT encoder with an autoregressive language backbone.
Meta's Chameleon explored early native multimodal tokenization by quantizing images into discrete tokens using a visual codebook.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Horizon Robotics model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.