Skip to main content
Speaker Diarization & Voice Extraction

AI Speaker Diarization Online

Upload a conversation, a scene or a podcast episode. CosyVoice speaker diarization separates who spoke when, clusters every voice into a distinct speaker, and turns each one into a named, reusable voice — no code, no GPU setup.

Diarization runs as a background job: vocal separation, transcription, then speaker clustering. A 10-minute file usually takes a few minutes — you can leave this page and come back.

How online speaker diarization works

Four diarization stages run in the cloud for every upload. Longer audio takes longer — expect minutes, not seconds, and results wait for you when the job finishes.

01

Vocal separation

Background music and effects are stripped first, so the diarization model hears clean speech — and the extracted voice references stay clean too.

02

Transcription

Speech recognition segments the audio with word timestamps, so every speaker turn arrives with its own transcript.

03

Speaker clustering

A voiceprint is computed for each segment; segments that sound alike cluster into one speaker. You never have to declare how many speakers to expect.

04

Voice references

Each detected speaker gets a clean reference clip assembled from their steadiest lines — ready to name, describe, and clone.

Why run speaker diarization online

No code, no GPU

pyannote and WhisperX are excellent diarization libraries — if you want Python, CUDA and tuning. Here speaker diarization is one upload in the browser.

Voices, not just labels

Most speaker diarization tools stop at “Speaker 0: 00:00–00:12”. CosyVoice also extracts each speaker’s actual voice so you can hear it and reuse it.

Name and describe every speaker

Turn Speaker 0, 1, 2 into named characters with voice notes — searchable by you and by the CosyVoice agent when it casts voices for new audio.

Built for long recordings

Interviews, episodes, table reads: diarization jobs queue in the background, report progress stage by stage, and keep results until you come back.

What separated voices unlock

Character voice cloning

Extract every character from a scene you have rights to, then synthesize new lines in each voice with CosyVoice voice cloning.

Podcast & interview notes

Identify speakers in audio and get per-speaker transcripts and speaking-time stats instead of one unlabeled wall of text.

Dubbing preparation

Split audio by speaker before dubbing so every voice in the scene keeps its own identity in the target language.

Voice datasets

Turn multi-speaker recordings into per-speaker corpora for TTS research, evaluation and fine-tuning.

Speaker diarization FAQ

What is speaker diarization?

Speaker diarization answers “who spoke when”: it segments an audio stream and groups the segments by speaker identity. CosyVoice runs speaker diarization online — upload a file and get labeled speakers with a timeline, transcripts and reusable voice references.

How do I identify speakers in an audio file?

To identify speakers in audio, upload the file and let the background diarization job finish. Each detected speaker comes back with a color-coded timeline, a speaking-time share and sample lines, so you can recognize and rename them in seconds.

Can I split audio by speaker into separate voices?

Yes — that's the core output. Every speaker gets an isolated, cleaned voice reference assembled from their best segments, which you can save as a cloneable voice in your library.

Can it extract voice from video, like a movie scene?

Yes. Upload video and the audio track is extracted automatically; vocal separation removes music and effects before speakers are clustered. Only upload material you own or have permission to use.

How is this different from pyannote or WhisperX?

Those are open-source diarization libraries for developers. This is a hosted pipeline: separation, transcription, speaker clustering and voice extraction run in the cloud, and the output plugs straight into voice cloning and the CosyVoice agent.

How long does speaker diarization take?

It isn't instant. The pipeline separates vocals, transcribes, and clusters voiceprints, so a 10-minute file typically takes a few minutes. Jobs run in a queue — close the tab and your results will be waiting.

How many speakers can it detect?

You don't have to specify a number. Diarization infers the speaker count from the audio itself, and works best with up to about eight distinct voices per recording.