billwuhao/Comfyui_HeyGem ? reverse-engineered prompt

Reverse engineered prompt

Build me a ComfyUI custom node for HeyGem digital humans that lets me turn a source video and audio into a talking avatar result. I want it to work like a simple node I can drop into a workflow, with an easy setup for local use and a Docker based option if needed. The output should support full body digital human video, keep the motion looking natural, and handle different output sizes without breaking.

Please make it practical for real use, so if the source clip is shorter than the audio, it can repeat or ping pong to fill the full length, and if the video is longer, it should automatically trim or capture the right part to match the audio duration. Keep the frame rate aligned between the input video and the synthesized output so the motion stays in sync. If you need to check current docs or compatibility details online, go ahead and do that.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab