NVIDIA Released Nemotron 3 Diarization Model
The 100-million parameter model enables real-time speaker identification in multi-party audio streams.
Updated on Sept. 23, 2026 in Artificial Intelligence

Live Poll
Do you trust AI models to accurately identify who is speaking in a recording?
NVIDIA has released the Nemotron 3 Diarization model, an open-weights system designed to track up to eight speakers simultaneously in real-time. The model is now available for commercial use under the OpenMDW License 1.1.
Why it matters
Automated speaker diarization—the process of attributing speech to specific individuals—is critical for transcribing and analyzing multi-party conversations. By providing a high-performance open model, NVIDIA is standardizing how systems process complex audio environments.
The model features 100 million parameters and achieved a 14.72% diarization error rate in benchmark tests. It was trained on 10,000 hours of real conversations and 82,611 hours of simulated audio mixtures.
The players
NVIDIA
A semiconductor and software firm specializing in GPU architectures and AI compute platforms.
The details
The system utilizes a 31-layer Transformer encoder—a neural network architecture that processes sequence data by weighing the importance of different parts—augmented with rotary positional embeddings to track temporal data. To support real-time streaming, it employs an Arrival-Order Speaker Cache and a FIFO (first-in, first-out) queue to maintain context across audio chunks. The model assigns audio channels based on signal arrival times and supports .wav, .flac, .opus, and .mp3 formats at 16 kHz.
Timeline
NVIDIA released the Nemotron 3 Diarization model on September 23, 2026.
The Tech Race
The release sets a new performance benchmark within the Voice Arena Diarization-Bench, establishing a competitive target for open-weights audio models. Researchers will now monitor how the model's 14.72% error rate holds up as the benchmarking suite completes its Version 1 evaluation.
Developers can download the model weights from Hugging Face for commercial applications provided they adhere to the OpenMDW License 1.1. Running the system requires a Linux environment equipped with NVIDIA GPUs spanning the Ampere to Blackwell architectures.
The takeaway
This model provides a scalable foundation for developers building transcription and analysis tools for multi-speaker environments. Watch for updated rankings in the Voice Arena Diarization-Bench to see if the 14.72% error rate remains the industry-leading benchmark.
Further reading
For more on the current state of acoustic modeling and speaker attribution, see the latest developments in Artificial Intelligence.
Live Poll
Do you trust AI models to accurately identify who is speaking in a recording?









