billythekidz/VibeVoice ? reverse-engineered prompt

Reverse engineered prompt

Build me a simple app for generating long conversational speech from text, like podcasts or dialogue scenes, using the VibeVoice model.

I want a page where I can paste a script, assign speaker names to different parts, and then generate an audio file I can listen to or download. It should handle single speaker and multi speaker scripts, and make the voices sound natural with turn taking, emotion, and the occasional background music or sound if the model does that.

Please include an easy demo interface, and a way to run it from a text file too. If it needs model weights or extra setup, make it clear in the app. Use whatever current docs or best practices you need to get it working, and make it stable enough that I can try both the smaller and larger model options without having to understand the internals.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab