- Sources: paper
- Summary: The preprint measures tokenization taking up to 64 percent of time to first token for coding agents running at a 94.1 percent prompt-cache hit rate. The measurement comes from 153,951 calls across two agent ecosystems, where the median call appends about 1.4K characters. The stated contract is that emitted token ids are always identical to full reference tokenization, and the paper reports a 16 to 34 percent drop in median time to first token under vLLM.
- Why it matters: Prompt-cache work moved the serving bottleneck onto a front end that was assumed cheap, and the identical-ids contract means the reported drop does not trade correctness for speed.
- Follow-up: Preprint with no peer review. Track whether the overhead figures reproduce outside the authors' two agent ecosystems.
send feedback on this story