deepspeedai/DeepSpeed ? reverse-engineered prompt

Reverse engineered prompt

Build me a working DeepSpeed project that helps me train and run a big PyTorch model efficiently on multiple GPUs. I want the library to make distributed training and inference feel simple, while still supporting the advanced stuff like memory saving, data and model parallel training, pipeline parallelism, and large model optimization. Please include a few clear example scripts so I can see how to plug it into a normal training loop and how to run inference with it.

I also want the setup to be practical for real use, so make sure it has the usual install and test flow, sensible docs, and a straightforward example for a large language model style workload. If there are current best practices or newer docs I should follow, look them up online if you need to. Keep the code organized so it would be easy for someone to try out on a GPU machine and understand what DeepSpeed is doing for them.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab