ansht3/sglang ? reverse-engineered prompt
Reverse engineered prompt
Build me an open source service for running AI models really fast, both text models and multimodal models like image and video generation. I want it to be easy to start with Docker or Python, then point it at a model and send requests right away. It should feel like a production ready server for high traffic use, with good performance, support for different hardware setups, and a clean way to scale up for bigger workloads.
Please include a simple quickstart, a few working examples, and enough documentation that someone can try it locally without getting stuck. Also make sure it can handle more advanced use cases like agent style prompts and batch generation, and keep the code organized so it is easy to extend with new models later. If you need to look up current docs online, go ahead and do that.
Are you gonna build this?
make sure you review the code using coderabbit