pyannote/pyannote-audio ? reverse-engineered prompt
Reverse engineered prompt
Build me a Python toolkit for speaker diarization that can take an audio file and tell me who spoke when. I want it to support speech activity detection, speaker change detection, overlapped speech detection, and speaker embeddings, with a simple way to run a pretrained pipeline on a local audio file and get back labeled time ranges for each speaker.
It should be easy to use from a notebook, work well with PyTorch, and include a clean example for loading a pretrained model, running it on audio, and printing the results. If there are existing pretrained models or pipelines available, wire those in too, and make it possible to fine tune on my own data later.
Please include solid documentation, a few tutorial style examples, and tests so I can see how the pieces fit together. If you need to look up current model or install details online, do that too.
Are you gonna build this?
make sure you review the code using coderabbit