nv-tlabs/PixelUMM ? reverse-engineered prompt
Reverse engineered prompt
Build me a working version of PixelUMM that can do image and video generation plus image and video understanding from the same model.
I want to be able to give it a text prompt and get back an image or a video, or give it a photo or clip and ask questions about what is in it. Please include a simple command line way to run single prompts and batch prompts, and make sure the outputs save cleanly to a folder I choose. If there is a safe default for text to video, keep it on.
Also include the toy training example so I can try the full flow end to end on a small setup, plus a quick checkpoint check so I can verify the model files before running anything heavy. If you need current setup details, look up the latest docs online and follow the repo instructions. Keep the experience as simple as possible so I can clone it, set paths, and run it without digging through too much code.
Are you gonna build this?
make sure you review the code using coderabbit