MehulMathodia/grpo-reasoner ? reverse-engineered prompt
Reverse engineered prompt
Build me a small Python project that fine tunes Qwen2.5 1.5B on GSM8K math problems with GRPO, using a simple verifiable reward based on whether the final boxed answer is correct. I want it to run on free or low cost compute, like a Kaggle P100, and include a clean baseline evaluation plus the trained adapter evaluation so I can compare before and after fairly on the full held out test set.
Please also add a basic supervised fine tuning option for comparison, a way to run maj at k style sampling checks, and a few plots that show training reward and benchmark results. Keep the reward and evaluation logic strict and reusable so they match exactly, and add tests for the answer extraction and scoring. If you need to check current library docs online, go ahead. I also want a simple demo and clear commands to train, evaluate, and reproduce the main results.
Are you gonna build this?
make sure you review the code using coderabbit