Latenode

NVIDIA releases open-weight audio model for 8-speaker diarization

Automation builders can add speaker-aware processing without relying entirely on a closed diarization API. Production workflows still need transcription, identity mapping and error checks around overlapping or noisy speech.

NVIDIA releases open-weight audio model for 8-speaker diarization

NVIDIA released Nemotron 3 Diarization, an open-weight audio model for tracking speakers in live and recorded conversations. The roughly 100-million-parameter checkpoint labels up to eight speakers, including overlapping speech, according to the launch post.

The model returns timestamps and anonymous speaker labels rather than identities or standalone transcripts. One checkpoint handles both streaming and offline-style processing, while weights are available on Hugging Face under OpenMDW 1.1.

Teams can pair its speaker channels with automatic speech recognition for meeting notes, call analysis and podcast processing. They will still need to map labels to known participants and test accuracy against noisy, domain-specific audio.

NVIDIA’s earlier Nemotron 3 release covers its language-model family; the diarization launch is a separate audio model.

When models and APIs shift keep the workflow running

Connect the models and apps from stories like this in one scenario. If a vendor changes routing, pricing, or availability, you update the workflow — you don't start over.

Start Free

Free forever plan. No credit card required.