Nvidia Releases Free 100M-Parameter Nemotron 3 Model for Speaker Diarization
- Published
- Sep 27, 2026 — 11:01 UTC
Nvidia has released the Nemotron 3 Diarization model, which features 100 million parameters and can identify up to eight speakers simultaneously. The model achieves a 14.7% diarization error rate (DER) on the VoiceArena Diarization Benchmark v1, outperforming its predecessor, Streaming Sortformer, by 41%. The Nemotron 3 model is designed to handle overlapping speech and minor misalignments, which are strictly scored as errors in the benchmark. In comparison, the next best system on the Diarization-Bench has a 19.3% error rate. The model operates with an audio buffer range from 30.4 seconds to 0.32 seconds, with a standard comparison buffer of 1.04 seconds. Jonathan Kemper noted that the weights of the model are freely available, making it accessible for developers and researchers looking to implement advanced speaker identification in their applications. This release follows Nvidia's ongoing efforts to enhance AI capabilities in speech recognition and diarization technologies.
By Callan Zhang · Sep 27, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: The Decoder
