NVIDIA/Megatron-LM ? reverse-engineered prompt

Reverse engineered prompt

Build me a GPU optimized training toolkit for huge transformer models that can scale across many GPUs and run distributed training smoothly.

I want it to include reusable building blocks for model layers, parallel training across tensor, pipeline, data, expert, and context modes, plus support for mixed precision so training can be faster and use less memory. Also include ready to run scripts for pretraining and a few example model setups, so someone can clone it and start experimenting without wiring everything together from scratch.

Make it easy to train, resume, and evaluate large language models, and add clear docs for setup, first training run, and how to use it on NVIDIA hardware. If there are current best practices or docs online that would help, check those too and follow them.

Are you gonna build this?

make sure you review the code using arcumet

Try freeSponsored — opens Arcumet in a new tab