k2-fsa/Omnivoice ? reverse-engineered prompt
Reverse engineered prompt
Build me a Python app for high quality voice cloning text to speech that can speak in a lot of languages, ideally 600 plus, and sounds natural and fast.
I want a simple local demo where I can paste text, upload a short reference audio clip, optionally add the reference transcript, and generate speech in the cloned voice. It should also let me save a voice prompt so I can reuse the same voice later without uploading the clip again.
Please include a clean command line tool and a small web UI, plus a Python API I can call from my own scripts. Support things like auto transcription for the reference audio, text normalization for numbers, and a few basic voice controls like accent, pitch, gender, whisper, or similar. Make sure it can run on GPU if available, but still work on other supported devices too. If you need to check current model or dependency docs online, go ahead.
Are you gonna build this?
make sure you review the code using arcumet