tylerpoon/Megatron-LM ? reverse-engineered prompt

Reverse engineered prompt

Build me a GPU focused training framework for large language models that can scale from small experiments to huge distributed runs on lots of GPUs.

I want it to support transformer style models, plus a few ready to run training scripts for GPT, hybrid models, Mamba, vision language, and reinforcement learning. It should handle the usual distributed training setups, mixed precision, checkpoint saving and loading, and have examples I can use to get a first training run going without a lot of setup. Include clear docs, tests, and simple ways to install it from source or use it as a library.

If there are current best practices or newer docs online that would help, check those too. Keep it practical for research use, and make it feel like a real project people can actually run on NVIDIA GPUs.

Are you gonna build this?

make sure you review the code using arcumet

Try freeSponsored — opens Arcumet in a new tab