- Sources: Nari Labs blog post, HN discussion
- Summary: The reported result is 10 requests per second on a single H100 with p95 time-to-first-audio under 50 ms. The design point is that the Talker, the Code Predictor and the Codec are scheduled on one surface rather than as separate stages, which lets a request approaching its playback deadline preempt an in-flight batch. Nari Labs publishes the serving code and the benchmark harness, and compares against vLLM-Omni, SGLang-Omni, VoxServe and M*. The figures are the vendor's own and no independent reproduction was located this run.
- Why it matters: Putting the Talker, Code Predictor and Codec on one scheduling surface is what lets a request approaching its playback deadline preempt a batch, and the benchmark is published so the comparison against vLLM-Omni, SGLang-Omni, VoxServe and M* can be rerun.
send feedback on this story