lucidrains/DALLE2-pytorch ? reverse-engineered prompt

Reverse engineered prompt

Build me a PyTorch project for DALL E 2 style text to image generation.

I want a clean library I can install, with simple scripts and examples for training the three main pieces, CLIP, the decoder, and the diffusion prior. The code should let me plug in my own text and image dataset, train on GPU, save checkpoints, and later use a trained model to generate images from text. Please include a few working sample configs and a small demo that shows how to run training and inference end to end.

If you need to check the current docs for anything, feel free to look them up online. Keep the API easy to understand, add basic README instructions, and make sure the example code actually runs without me having to guess missing steps.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab