• Sources: GigaToken repo, HN 49010167
  • Summary: GigaToken is an MIT-licensed Rust tokenizer with Python bindings that reports GB/s throughput. On a 144-core AMD EPYC it measured 989x over Hugging Face tokenizers and 681x over tiktoken for GPT-2 BPE on the 11.9 GB OpenWebText set, with smaller gains for SentencePiece models. The approach uses SIMD pretokenization in place of a regex engine, caches pretoken mappings for long-tailed words, and minimizes Python-to-Rust overhead. Benchmarks span AMD EPYC, Apple M4 Max, and Ryzen, and the authors note GigaToken encodes whole files while the baselines process pre-split samples.
  • Comments: HN commenters ask which real workloads are tokenizer-bound, since tokenization rarely dominates training or inference wall-clock, while praising the SIMD engineering.
  • Why it matters: It removes tokenization as a preprocessing bottleneck for large corpora, though the practical benefit depends on a pipeline actually being tokenizer-bound.

send feedback on this story