DLR-RM/stable-baselines3 ? reverse-engineered prompt
Reverse engineered prompt
Build me a Python library for reinforcement learning in PyTorch that feels easy to use but still solid and reliable.
I want it to let me train common RL agents on gym style environments with a simple, sklearn like API, and support things like custom environments, custom policies, callbacks, TensorBoard logging, and dict observations. It should be friendly for notebooks too, and include clear docs and examples that show how to train, save, load, and run models on a basic task like CartPole. Please make the code clean, well tested, and easy to extend, since this is meant to be a stable core people can build on top of.
If you need to, look up current docs online so the setup and examples match modern versions of the libraries involved.
Are you gonna build this?
make sure you review the code using coderabbit