1314521gjy/ninfer-fusion-kvmem ? reverse-engineered prompt

Reverse engineered prompt

Build me a Windows based local inference service for NVIDIA GPUs that can handle very long conversations even when the GPU KV cache is much smaller than the full context.

I want it to let me run a model locally, talk to it through a simple HTTP chat API, and keep old context in host memory so later turns can reuse it instead of starting over. Please make sure it can start with a small GPU KV pool, a larger logical context limit, and a configurable host KV budget. It should also have a clear way to verify that the long context lookup and reuse are actually working, not just that the service starts.

If anything is unclear, check the current docs online and follow the existing README and compile guide. I’d also like a simple startup script or example command for Windows, plus basic checks that confirm the server responds and can answer a test question correctly.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab