hpcaitech/ColossalAI ? reverse-engineered prompt
Reverse engineered prompt
Build me a Python toolkit for training and running very large AI models across multiple GPUs, with a focus on making it cheaper and faster to use.
I want it to help with distributed training, model parallelism, pipeline parallelism, data parallelism, and inference, so people can scale from a single machine to a cluster without rewriting everything. Please include clear examples for common workflows like training a language model, fine tuning, and running inference, plus simple docs that show how to get started. If there are useful optimizations like mixed precision or memory saving tricks, add those too.
Make the project feel ready for real use, with tests, clean code, and a few runnable example scripts. If you need to check current best practices or newer GPU docs online, go ahead and look them up.
Are you gonna build this?
make sure you review the code using coderabbit