Skip to content
HChuggingface.co·

tokenizers v1: encode, decode and scaling, measured

AI summary

Tokenizers v1 significantly improves performance, encoding text 3 to 30 times faster than v0.23 on an Apple M4 Max with a single thread, depending on the model family (e.g., t5-base to gpt2). It also demonstrates strong scalability, achieving 76% of linear scaling across eight workers. Crucially, v1 maintains exact token ID consistency with the previously released library, ensuring no changes in output despite the performance enhancements.

Why this one

Unlike prior versions, tokenizers v1 delivers a 3 to 30 times speed improvement over v0.23 while maintaining identical token IDs, moving tokenization from a potential bottleneck to a performance accelerator.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 21, 2026, 00:00 UTC

IngestedOffset at this time: UTC+0Sep 21, 2026, 15:02 UTC

Published
Sep 21, 2026, 00:00
Ingested
Sep 21, 2026, 15:02
Source type
Official
Tier
First-party
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

Discussion trend

→ Steady
Latest 24h versus previous 24h snapshot means · 7-day curve

The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.

The tokenizer has not historically been the bottleneck within ML workflows. Compute-wise, tokenization is light compared to the heavy modeling happening in the rest of the pipeline. Yet, in some cases, it has rapidly become key to accelerating (or slowing down) your machine learning work.

As models become faster and workloads scale, that balance begins to shift. Training on massive datasets, serving many concurrent requests, or repeatedly processing long inputs can put enough pressure on the tokenizer that it starves the model of data.

This is why we have chosen to heavily focus on performance for the upcoming version 1 of tokenizers. Tokenization should be light and should scale with your workflow. Your GPUs should never sit idle waiting for the CPU to complete its tokenization.

In this article, we look at what makes v1 faster than v0.23, often by tens of times.

This work was entirely possible thanks to the rest of the ecosystem. Tokenization is a very active area of open source work, and libraries such as gigatoken, tiktoken, kitoken, tokie, fastokens, wordchipper and ai-tokenizer, as well as many others, have each pushed on what a fast tokenizer can be. We read that work, and several of the ideas below reached us because another project showed they were worth trying.

Before this refactor, tokenizers was nowhere near the performance it could have had, so contributing to it may not have seemed worth it. With this refactor, we hope to make clear that we intend tokenizers to be a library worth contributing to.

We also thank IBM, NVIDIA, and the ExecuTorch team for contributing patches and helping us test across a wide range of hardware to broaden platform support.

Results

We showcase results for the release candidate of tokenizers v1 against other widely used alternatives. We go over single-threaded, multi-threaded, scaling across threads, per-model comparison, per-language comparison, latency, decoding throughput, memory heap, as well as crate size.

We run this from the tokbench repository, and add a command to rerun the benchmarks on your hardware if you would like to do so.

What V1 Is

v1 will produce the same token IDs as v0.23. The goal was to preserve the output, the API, the vocabulary and the merge ranks, and improve everything that can be improved. That includes breadth. The library stays general across tokenizer families rather than specialising on BPE, so v1 loads everything v0.23 loaded.

A tokenizer converts text into the list of integers a model reads. tokenizers runs that conversion in four stages. Normalization applies operations such as lowercasing or Unicode normalization to the raw text. Pre-tokenization splits the text into smaller pieces called pre-tokens. The model turns each pre-token into tokens and maps them to IDs in its vocabulary. Post-processing adds any special tokens the model expects.

The model stage is where most of the work described here happens. Eight of the ten model families measured in this article use byte pair encoding, or BPE. BPE starts from the bytes of a pre-token and repeatedly joins the highest ranked adjacent pair until no ranked pair remains. The ranking is learned when the tokenizer is trained and ships with it, so the same text always produces the same IDs. A merge never crosses a pre-token boundary. The other two families use WordPiece and Unigram, the two other model types the library supports.

The tokenization pipeline page documents the four stages. Tokenization algorithms documents BPE, WordPiece and Unigram.

Each stage was worked on. These are the changes that mattered:

change what it does

workspace split one crate became a workspace: tk-encode is the required runtime, and tk-serialize, tk-convert and tk-train are linked only when an application needs them

no-alloc model the merge working set lives in a caller-owned scratch buffer; the loop never touches the allocator

bitcannon the split pattern becomes Boolean operations over bitstreams, using SIMD instructions to find splits instead of a regex engine

merge-loop rewrite the pieces being merged form an intrusive doubly-linked list inside one preallocated buffer, so a merge updates two indices instead of moving data

word cache a thread-local memo from pre-token bytes to finished ids, so a repeated word is merged once

native parallelism one shared tokenizer encodes from many threads at once; each thread draws its scratch buffer and word cache from its own sub-pool, so threads no longer queue on a single lock (#2365)

The Split: Bitstreams Instead Of A Regex