brohan203/freetoken-vulkan ? reverse-engineered prompt
Reverse engineered prompt
Build me a small app that can run OpenAI gpt oss 20b on an AMD Radeon GPU through Vulkan, with no CUDA, ROCm, or Triton.
I want it to load the model weights from a local folder, build the shaders, and let me chat with it from the command line or a simple demo script. It should support long lived generation so repeated prompts get faster, and it should keep the GPU memory under control instead of leaking over time. If possible, include support for the smaller dense Qwen3 model too, so I can compare speed and behavior on the same machine.
Please make the whole thing work on Windows first, with clear setup steps for building the shaders and running a test prompt. Use current docs online if you need to. I also want basic correctness checks, so I can confirm the outputs are reasonable and the model is stable for a few hundred tokens.
Are you gonna build this?
make sure you review the code using coderabbit