Ai2 released new Bolmo byte-level language model checkpoints on Hugging Face on October 7, 2026. Bolmo models operate directly on UTF-8 bytes instead of subword tokens, which improves character-level understanding and multilingual text handling. The release extends the approach beyond the original Olmo-based models to include Bwen 8B and Llama-B 8B, along with new Stage 1 checkpoints. The models are open and licensed under Apache 2.0.
A situational dense language model from Ai2. Sits below the open-model average on our benchmarks; pick it for specific deployment constraints rather than peak quality. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 0.5 GB | Low | |
| Q4_K_MRecommended | 0.5 GB | Good | |
| Q5_K_M | 0.5 GB | Very Good | |
| Q6_K | 0.5 GB | Excellent | |
| Q8_0 | 0.5 GB | Near Perfect | |
| FP16 | 0.5 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| AMD Radeon RX 7600 8GBAMD | AA | 463.7 tok/s | 0.5 GB |
| GEEKOM IT15 (Ultra 9 285H)GEEKOM | AA | 144.9 tok/s | 0.5 GB |
| Lenovo ThinkCentre P3 Tiny Gen 2 (Ultra 5 235)Lenovo | AA | 144.9 tok/s | 0.5 GB |
| NVIDIA GeForce RTX 4060NVIDIA | AA | 437.9 tok/s | 0.5 GB |
| NVIDIA GeForce RTX 5060 Ti 8GBNVIDIA | AA | 721.3 tok/s | 0.5 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.07 |
NVIDIA GeForce RTX 4060Vast.ai · On-Demand · 8 GB VRAM | $0.08 |
NVIDIA GeForce RTX 3070Vast.ai · Spot · 8 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Bolmo is an open byte-level language model family developed by the Allen Institute for AI (Ai2) that replaces traditional subword tokenization with direct UTF-8 byte processing. Released under the Apache 2.0 license, Bolmo addresses the fundamental design flaws of subword tokenizers (such as Byte-Pair Encoding and WordPiece), including vocabulary bloat, poor character manipulation, brittle whitespace handling, and structural bias against non-Latin scripts. On October 7, 2026, Ai2 expanded the ecosystem on Hugging Face by publishing new checkpoints, including retrofitted versions of popular open backbones such as Bwen 8B and Llama-B 8B alongside intermediate Stage 1 checkpoints.
Rather than training a byte-level transformer from scratch, which has historically demanded prohibitive compute budgets, Ai2 retrofits existing pretrained models into byte-level systems using less than 1% of the original pretraining compute. The model uses a dense architecture with undisclosed parameter distributions across its internal layers and byte-conversion modules. By stripping away hand-engineered token vocabularies, Bolmo provides developers with a robust local foundation model that excels at character-level manipulation, structured data parsing, and multilingual generation without tokenization artifacts.
For practitioners looking to deploy a local AI model undisclosed parameters 2026 release, Bolmo occupies a unique position. It preserves the general reasoning capabilities of mainstream subword models while eliminating edge cases where tokenizers fragment words, garble source code formatting, or bloat compute when handling non-English text.
Bolmo departs from standard causal transformer designs by removing the subword embedding table and vocabulary projection head. Standard language models project hidden states across vocabularies spanning 32,000 to 128,000 discrete token IDs. In contrast, Bolmo operates over an alphabet of 256 raw UTF-8 bytes plus a minimal set of control bytes.
Ai2 constructs Bolmo via a two-stage retrofitting process:
Because byte sequences are substantially longer than subword token sequences (often requiring 3x to 4x the sequence length to represent the same text), raw full attention over bytes is computationally prohibitive. To mitigate this, Bolmo integrates localized recurrent modules using extended Long Short-Term Memory (mLSTM / xLSTM) blocks alongside optimized flash attention. These local modules condense incoming byte streams into grouped representations before feeding them through the dense transformer core.
Although the exact context length is not specified in the initial base release, the architecture is designed to handle extended input sequences through efficient local byte compression, ensuring that sequence expansion does not exhaust VRAM during practical inference.
Bolmo is built as a text-only system tuned for code execution, rigorous reasoning, and cross-lingual robustness.
Subword models struggle with character counts, spelling tasks, anagrams, and string transformations because individual letters are swallowed into unified token IDs. Bolmo processes text byte-by-byte, making it naturally immune to these blind spots. On any standard Bolmo reasoning benchmark evaluating exact string matching, cryptographic hashing logic, or regular expression parsing, the model avoids the classic failure modes of subword networks.
When deploying Bolmo for coding, the byte-level architecture offers distinct advantages over BPE-based models:
Standard tokenizers heavily penalize non-Latin alphabets. In languages like Arabic, Hindi, Japanese, or Cyrillic, a single word may be split into four to eight tokens, driving up latency and inference cost. Bolmo standardizes the compute cost across languages by evaluating raw UTF-8 bytes, creating a more uniform distribution of compute across global languages.
To run Bolmo locally, your hardware must accommodate both the dense transformer backbone and the local byte-processing modules (which rely on the xlstm library).
Because byte sequences require smaller vocabulary heads (256 classes instead of 128,000), memory consumption at the final classification layer is minimal. However, you must account for activation memory caused by longer raw byte sequences.
Below are the hardware requirements based on typical dense deployments in 8B-class configurations:
The best quantization for Bolmo is Q4_K_M or standard AWQ/GPTQ 4-bit for consumer setups. Because the model relies on a dense transformer backbone with byte-level adapters, 4-bit weight quantization preserves core reasoning and coding ability while slashing the base memory footprint under 8 GB. If your workload involves dense mathematical reasoning or character-critical cipher evaluation, step up to Q8_0 to preserve the precision of the xLSTM byte projection layers.
When analyzing Bolmo performance, raw metrics differ from standard models. In a subword system, generation speed is measured in subword tokens per second. With Bolmo, inference metrics are tracked in raw bytes per second:
To run Bolmo on your local machine using Python, install PyTorch along with transformers (version 4.57.3 or later) and xlstm:
1pip install "transformers>=4.57.3" "xlstm>=2.0.4"
You can initialize inference using Hugging Face's AutoModelForCausalLM:
1from transformers import AutoModelForCausalLM, AutoTokenizer2import torch34device = "cuda" if torch.cuda.is_available() else "cpu"5model_id = "allenai/Bolmo-7B"67tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)8model = AutoModelForCausalLM.from_pretrained(9 model_id,10 torch_dtype=torch.bfloat16,11 device_map="auto",12 trust_remote_code=True13)1415inputs = tokenizer(["def reverse_string(s: str) -> str:"], return_tensors="pt").to(device)16# Note: max_new_tokens controls the generated byte length17output = model.generate(**inputs, max_new_tokens=128, temperature=0.1)18print(tokenizer.decode(output[0], skip_special_tokens=True))
If you are investigating how to run undisclosed model on consumer GPU setups using high-level managers, Ollama support for Bolmo relies on backend forks incorporating custom byte tokenizers and xLSTM operations. While community GGUF ports are rolling out, native PyTorch execution via the official Ai2 bolmo-core repository remains the most stable method for local inference and fine-tuning.
Understanding when to run Bolmo versus established open-weight alternatives comes down to architectural tradeoffs:
Meta's Llama 3 8B utilizes a massive 128,000-token Tiktoken vocabulary. In standard conversational English, Llama 3 generates text faster per word because each forward pass produces multi-character chunks. However, Llama 3 suffers from standard tokenization traps: spelling errors on unfamiliar acronyms, vulnerability to token manipulation exploits, and fragmented generation on code featuring heavy whitespace or atypical syntax. Bolmo matches Llama 3's high-level reasoning while completely sidestepping tokenization errors, making it superior for syntax parsing, hex data processing, and precise text modification.
Alibaba's Qwen 2.5 7B is known for its strong multilingual and coding capabilities. However, its subword tokenizer remains biased toward common Asian and Western languages, degrading efficiency when faced with low-resource dialects or mixed-script strings. Bolmo processes all inputs as raw UTF-8 bytes, providing an even playing field across every human script without requiring regional vocabulary expansion. While Qwen holds a slight speed advantage on conventional benchmarks, Bolmo is more predictable on unusual, character-critical workloads.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Allen Institute for AI model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.