QwenAudio/CosyVoice ? reverse-engineered prompt

Reverse engineered prompt

Build me a simple app for multilingual text to speech that lets me paste in text, pick a voice or upload a short reference clip, and generate natural sounding audio. I want it to handle Chinese, English, Japanese, Korean, and a few other common languages, plus support voice cloning and different speaking styles like emotion, speed, and volume.

Please include a clean web page where I can type text, hear the result, and download the audio. It would be great if it also supports streaming playback for longer text so I do not have to wait for everything to finish. If needed, look up the current model and runtime docs online and wire it up in the easiest way that works with the repo. Make sure the default experience is simple enough for someone who just wants to try it locally, but keep the code organized so it could also be used in a demo or API later.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab