Youtu-Parsing-Omni is a 5.33B open-weight omni-modal parsing model from Tencent's Youtu team. It takes a document page, natural image, chart, flowchart, geometry figure, audio clip or audio-visual video and returns one structured JSON object covering layout elements, text, LaTeX, tables, bounding boxes, timestamps, OCR, ASR and captions. The model card reports 96.96 overall on OmniDocBench v1.6 and 75.08 average on OmniParsingBench. Weights, a vLLM plugin and inference examples are public under a custom license that states the model is not intended for use within the European Union.
A workable 5.33B-parameter dense language model from tencent. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 5.3 GB | Low | |
| Q4_K_MRecommended | 6.4 GB | Good | |
| Q5_K_M | 6.9 GB | Very Good | |
| Q6_K | 7.6 GB | Excellent | |
| Q8_0 | 8.9 GB | Near Perfect | |
| FP16 | 14.0 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| AMD Radeon RX 7700 XTAMD | SS | 54.5 tok/s | 6.4 GB |
| Intel Arc B580Intel | SS | 57.5 tok/s | 6.4 GB |
| NVIDIA GeForce RTX 4070NVIDIA | SS | 63.5 tok/s | 6.4 GB |
| NVIDIA GeForce RTX 4070 SUPERNVIDIA | SS | 63.5 tok/s | 6.4 GB |
| NVIDIA GeForce RTX 5060 Ti 8GBNVIDIA | SS | 56.5 tok/s | 6.4 GB |
Energy cost on Raspberry Pi 5 (8GB) (~4.3 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Youtu-Parsing-Omni on Raspberry Pi 5 (8GB) · ~4.3 tok/s · 12W | $0.093 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 6 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.07 |
NVIDIA GeForce RTX 4060Vast.ai · On-Demand · 8 GB VRAM | $0.08 |
NVIDIA Tesla V100 16GBVast.ai · Spot · 16 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Youtu-Parsing-Omni is a 5.33B parameter dense multimodal model developed by Tencent's Youtu team, designed to replace fragmented extraction pipelines with a single end-to-end parser. Instead of routing document images to optical character recognition (OCR) engines, layout analyzers, formula extractors, speech-to-text models, and video description pipelines separately, Youtu-Parsing-Omni processes document pages, charts, flowcharts, geometric diagrams, audio clips, and audio-visual video files through one pipeline. It returns a single, strictly formatted JSON payload containing bounding boxes, reading order, LaTeX formulas, table formats, timestamps, and captions.
The model marks a shift toward unified structural extraction models in the local AI ecosystem. While general-purpose vision-language models can describe images or transcribe text when prompted, they frequently fail at structural consistency, coordinate precision, and strict JSON output formatting. Youtu-Parsing-Omni directly targets these enterprise ingestion bottlenecks, posting a 96.96 overall score on OmniDocBench v1.6 and a 75.08 average on OmniParsingBench.
For developers deploying on local infrastructure, the 5.33B dense architecture hits a practical sweet spot between capability and resource overhead. It allows teams running dedicated workstations or edge servers to bypass cloud extraction APIs entirely. However, practitioners must evaluate the model's custom open-weight license before deployment, as it explicitly contains terms stating the weights are not intended for use within the European Union.
Youtu-Parsing-Omni uses a dense 5.33B parameter architecture optimized for high-throughput perception and serialization tasks. Building upon Tencent's earlier Youtu-LLM and Youtu-Parsing systems, the omni variant integrates a dynamic visual encoder capable of preserving fine-grained document features alongside unified audio and temporal encoders.
A critical technical specification is its native 65,536-token context length. Standard multimodal models often cap context windows at 4,096 or 8,192 tokens, which creates immediate bottlenecks when outputting dense OCR, hierarchical layout bounding boxes, and LaTeX equations for high-resolution documents or lengthy audio-visual tracks. A single complex financial sheet or technical drawing can easily generate tens of thousands of characters of structured output. With a 64k-token window, Youtu-Parsing-Omni handles extended document pages, continuous audio recordings, and video sequences without truncating output structures mid-generation.
Because the model is fully dense, all 5.33B parameters participate in every inference pass. This creates consistent, predictable memory bandwidth saturation and predictable generation latency, differing from Mixture-of-Experts (MoE) architectures that swap active parameter blocks dynamically. Tencent provides integration via official inference scripts and a dedicated vLLM plugin, allowing developers to take advantage of continuous batching and PagedAttention to maximize throughput on local hardware.
Youtu-Parsing-Omni combines low-level sensory perception with semantic interpretation. Its primary output format across all modalities is structured JSON, conditioned by explicit task prompts.
bbox) for every element, calculates logical reading orders across multi-column formats, parses tabular data into HTML or OTSL (Open Table Structure Language), and translates mathematical notation into raw LaTeX.For engineering teams building local retrieval-augmented generation (RAG) engines, this model replaces complex multi-stage ingestion stacks (like running Tesseract, layoutlm, Whisper, and a chart extraction script simultaneously) with a single unified execution graph.
The 5.33B parameter footprint makes local hardware requirements accessible on consumer and prosumer equipment. Memory consumption is dictated by the model weights, the vision/audio encoder projections, and the large 65,536-token KV cache when processing dense outputs.
| Quantization Level | Weight Footprint | Minimum VRAM (Short Context / 4k) | Recommended VRAM (Full 64k Context) |
|---|---|---|---|
| FP16 / BF16 | ~10.7 GB | 14 GB | 24 GB |
| Q8_0 | ~5.8 GB | 9 GB | 16 GB |
| Q4_K_M | ~3.3 GB | 6 GB | 12 GB |
When calculating how to run a 5.33B model on consumer GPU setups, remember that high-resolution visual tokens consume noticeable activation memory. For users seeking the best quantization for Youtu-Parsing-Omni, a 4-bit format like Q4_K_M or AWQ 4-bit preserves structural JSON syntax without degrading bounding box accuracy, allowing the entire model to run within modest memory ceilings.
On an RTX 4090 running Tencent's vLLM plugin, Youtu-Parsing-Omni tokens per second typically hover between 65 to 90 tokens/sec for plain text generation phases, slowing during the initial multimodal encoding step depending on image resolution or video frame counts. On a single RTX 3060 running 4-bit quantization, expect sustained generation rates of 30 to 45 tokens/sec.
To run Youtu-Parsing-Omni locally, developers can deploy the code provided in the official TencentCloudADP/youtu-parsing repository using Python 3.10+ and CUDA 12.x. While Ollama integration depends on community conversions to GGUF format, the native vLLM engine currently provides the lowest latency and most stable token-parallel decoding.
Evaluating Youtu-Parsing-Omni vs alternative local vision models clarifies where it fits in an engineering pipeline:
For developers seeking a capable local AI model with 5.33B parameters in 2026, Youtu-Parsing-Omni delivers a production-grade multi-modal ingestion pipeline on accessible consumer GPU hardware, provided your deployment falls outside the restrictions of its custom license.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Tencent model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.