• Sources: Hugging Face blog post, discussion
  • Summary: The post reports the model ranking first among 12 systems and 17 configurations on VoiceArena's Diarization-Bench, with a diarization error rate of 14.72% against 19.3% for the next-ranked system, and a 41.0% unweighted mean relative error reduction over the diar_streaming_sortformer_4spk-v2.1 baseline at 1.04 seconds of latency. The model is 100M parameters and handles up to eight speakers. It lists four recommended input-buffer operating points, 30.4, 1.04, 0.64 and 0.32 seconds, and notes 0.32 seconds is the lowest recommended setting rather than a floor, covering offline batch and streaming operation from one checkpoint. The source is a Hugging Face community blog post authored by NVIDIA staff rather than a model card, published 2026-09-23, so every figure is Nvidia's own.
  • Why it matters: One open-weight 100M-parameter model reported at the top of a public diarization benchmark makes speaker-attributed transcription practical to self-host instead of buying it as an API.

send feedback on this story