NVIDIA Released Nemotron 3 Diarization Model

The 100-million parameter model enables real-time speaker identification in multi-party audio streams.

Updated on Sept. 23, 2026 in Artificial Intelligence

Isometric editorial illustration featuring a brass audio connector cable snaking through clean geometric matte forms, representing real-time audio processing technology.
NVIDIA has released the Nemotron 3 Diarization model, an open-weights system designed to identify and track up to eight speakers simultaneously in real-time. AI Illustration. Upload story photo >

Live Poll

Do you trust AI models to accurately identify who is speaking in a recording?

NVIDIA has released the Nemotron 3 Diarization model, an open-weights system designed to track up to eight speakers simultaneously in real-time. The model is now available for commercial use under the OpenMDW License 1.1.

Why it matters

Automated speaker diarization—the process of attributing speech to specific individuals—is critical for transcribing and analyzing multi-party conversations. By providing a high-performance open model, NVIDIA is standardizing how systems process complex audio environments.

The model features 100 million parameters and achieved a 14.72% diarization error rate in benchmark tests. It was trained on 10,000 hours of real conversations and 82,611 hours of simulated audio mixtures.

The players

NVIDIA

A semiconductor and software firm specializing in GPU architectures and AI compute platforms.

The details

The system utilizes a 31-layer Transformer encoder—a neural network architecture that processes sequence data by weighing the importance of different parts—augmented with rotary positional embeddings to track temporal data. To support real-time streaming, it employs an Arrival-Order Speaker Cache and a FIFO (first-in, first-out) queue to maintain context across audio chunks. The model assigns audio channels based on signal arrival times and supports .wav, .flac, .opus, and .mp3 formats at 16 kHz.

Timeline

  1. NVIDIA released the Nemotron 3 Diarization model on September 23, 2026.

The Tech Race

The release sets a new performance benchmark within the Voice Arena Diarization-Bench, establishing a competitive target for open-weights audio models. Researchers will now monitor how the model's 14.72% error rate holds up as the benchmarking suite completes its Version 1 evaluation.

Developers can download the model weights from Hugging Face for commercial applications provided they adhere to the OpenMDW License 1.1. Running the system requires a Linux environment equipped with NVIDIA GPUs spanning the Ampere to Blackwell architectures.

The takeaway

This model provides a scalable foundation for developers building transcription and analysis tools for multi-speaker environments. Watch for updated rankings in the Voice Arena Diarization-Bench to see if the 14.72% error rate remains the industry-leading benchmark.

Further reading

For more on the current state of acoustic modeling and speaker attribution, see the latest developments in Artificial Intelligence.

Live Poll

Do you trust AI models to accurately identify who is speaking in a recording?