NVIDIA released Nemotron 3 Diarization, an open-weight audio model for tracking speakers in live and recorded conversations. The roughly 100-million-parameter checkpoint labels up to eight speakers, including overlapping speech, according to the launch post.
The model returns timestamps and anonymous speaker labels rather than identities or standalone transcripts. One checkpoint handles both streaming and offline-style processing, while weights are available on Hugging Face under OpenMDW 1.1.
Teams can pair its speaker channels with automatic speech recognition for meeting notes, call analysis and podcast processing. They will still need to map labels to known participants and test accuracy against noisy, domain-specific audio.
NVIDIA’s earlier Nemotron 3 release covers its language-model family; the diarization launch is a separate audio model.
