MohammadHaishemKhawaja/hashyy ? reverse-engineered prompt

Reverse engineered prompt

Build me a simple Windows friendly setup for running this local llama.cpp based model server on one 12 GB GPU, with one launch option for the big 177B model and another for the faster 35B mode.

I want double click batch files that start the server on port 8080, wait until it is ready, and then let me open it in a browser or connect to it from tools that use an OpenAI style API. It should support chat completions and tool calling, and it should be easy to switch between the big model and the faster profile without editing code.

Please keep the benchmark and quality checking pieces in place too, so I can measure speed and compare output against a control run. If anything needs current docs or build steps, look them up online and make the scripts as reliable as possible. Also make the startup and troubleshooting messages clear for a normal person, since I am mainly trying to run this locally, not study the internals.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab