How to Use the WhisperX Diarization Pipeline Prompt to Get Speaker Labeled Transcripts
Plan a WhisperX transcription run with speaker labels: faster-whisper model size and compute type for your GPU, batch size, language and word alignment, pyannote speaker-diarization-3.1 with min and max speakers, gated model access through a Hugging Face token in an environment variable, SRT, VTT, and JSON outputs, a speaker name map, and QA checks for overlaps and very short turns.

Getting a transcript out of Whisper is easy. Getting a transcript where every line is labeled with the right speaker, the timestamps line up with the audio, and the run does not crash on hour three of a batch is harder. WhisperX combines a fast Whisper backend, word level alignment, and pyannote speaker diarization, but the settings that matter are spread across model size, compute type, batch size, gated model access, and output formats. The Speaker Diarized Transcription Pipeline Planner With WhisperX and pyannote speaker-diarization-3.1: Model Size vs VRAM, compute_type, batch_size, Alignment, Speaker Counts, HF Token via Env Var, SRT VTT JSON Outputs, and QA prompt plans the whole pipeline for your hardware and audio, then adds QA checks so you catch bad speaker labels before they reach an editor.
What the prompt produces
- A model plan choosing a faster-whisper model size and compute_type for your GPU or CPU, with a step to measure peak memory instead of trusting guesses.
- batch_size guidance, including the rule to lower batch size before shrinking the model when memory runs out.
- Language and alignment settings, with the language set explicitly and a word alignment model for that language.
- Diarization setup using pyannote speaker-diarization-3.1, min and max speakers, and the steps to get access to the gated models with a Hugging Face token kept in an environment variable.
- A CLI command and a Python batch script that skips files already processed.
- Output plans for SRT, VTT, and JSON with word timings and speaker labels.
- A speaker name map that turns generic labels into real names per file.
- QA checks for words with no speaker, very short turns, overlapping speech, unexpected speaker counts, and timestamp drift.
- A troubleshooting table for the errors people hit most.
How to fill the inputs
AudioSet describes the files: format, total hours, typical length, number of speakers, mic setup, crosstalk, and music.
Hardware lists your GPU and its memory, or CPU only, plus OS, Python, CUDA, and the WhisperX version you installed. Function names change between releases, so the version matters.
LanguageProfile gives the language, accents, and names or technical terms likely to appear.
SpeakerInfo sets the expected range of speakers and who they are.
OutputSpec lists the formats you need and who uses them, such as captions for an editor or JSON for search.
Priority says whether accuracy or speed matters more, and the deadline.
Reading the example output
The example plans a run for 60 podcast episodes on a 12 GB consumer GPU:
- large-v3 with float16 is the starting choice, with an instruction to check peak memory on the first episode and switch to a lower precision compute type if needed.
- Access setup is a one time step: accept the conditions on both gated pyannote model pages and export a read token as HF_TOKEN.
- The CLI and Python versions match, and the Python script writes a done flag per episode so a crash does not mean starting over.
- Import paths are marked CHECK, since the diarization class moved between WhisperX versions.
- The speaker map is stored per episode, because SPEAKER_00 in one file is not necessarily the same person in the next.
- QA checks have a fix for each problem, such as setting min speakers to 2 when everything comes out as one speaker.
Tips for better results
- Run one representative file end to end before starting a batch.
- Trim long music intros, which can produce invented text.
- Set min and max speakers when you know the count. It often improves labels.
- Spot check a few minutes of every file with the audio playing, especially around crosstalk.
Mistakes to avoid
- Do not paste a Hugging Face token into scripts or notebooks you share.
- Do not rely on auto language detection for files that start with music or silence.
- Do not assume speaker labels carry across files.
- Do not skip alignment if you need word timings for captions or search.
Who it is for
Podcast producers, researchers transcribing interviews, journalists working with recorded calls, and engineers building transcription pipelines for meetings or archives.
Related PromptDig links
Open the Speaker Diarized Transcription Pipeline Planner With WhisperX and pyannote speaker-diarization-3.1: Model Size vs VRAM, compute_type, batch_size, Alignment, Speaker Counts, HF Token via Env Var, SRT VTT JSON Outputs, and QA prompt and describe your audio and hardware. For more AI tools prompts, Browse more prompts. If you have a local AI workflow prompt that works, Share a prompt.