lloyal-ai/liblloyal ? reverse-engineered prompt

Reverse engineered prompt

Build me a C++20 library that acts like a little branching engine for live model inference state.

I want to be able to create a branch, fork it, prune branches, keep only one winner, and reuse shared prefix state so I do not recompute tokens that every branch already has. It should feel a bit like Git for generation trees, where each branch can have its own sampler, seed, grammar, logits, and metrics, but still share the same history until they diverge.

Please also make it support advancing many branches together in one pass, handling both single token and multi token updates, plus a way to blend logits from multiple branches without copying state. Images should be able to come in as embeddings too, so multimodal inputs can live in the same tree.

Keep it clean, documented, and testable, with a small public API and a working example that shows fork, decode, prune, and retain only one branch. If you need current llama.cpp details, look up the latest docs online.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab