gravitee-io/gravitee-singularitee ? reverse-engineered prompt

Reverse engineered prompt

Build me a production ready inference server for LLM pipelines in Java.

I want a service that can load a workspace YAML, download model weights on first start, and serve models, classifiers, embeddings, reranking, and multi step pipelines from one place. It should expose gRPC as the main API, and also an OpenAI compatible HTTP API for chat completions plus endpoints for classification, reranking, and similarity. Streaming should be supported end to end, and the server should be able to run models through different backends like llama.cpp, vLLM, ONNX, and GLiNER depending on what the workspace asks for.

Make it easy to run locally with a simple install and start command, and include a few example workspaces and smoke tests. If helpful, look up current docs online for the model and backend libraries before wiring things up.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab