dataflowr/llm_efficiency ? reverse-engineered prompt

Reverse engineered prompt

Build out this homework project for me, using the existing minGPT code as the base.

I want the first part to add KV cache support so generation does not recompute the whole prompt every time. It should still produce the same results as the normal generate method, but be faster during autoregressive decoding. Please make the sort demo work and include the benchmark that compares cached generation against the baseline.

Then do the LoRA part too. I want a small trainable adapter layer that freezes the original weights, learns only the low rank update, and can be merged back in for inference with no extra overhead. Wire it into the model where needed, finish the sorting fine tuning demo, and make sure it handles longer sequences better after training.

Please keep the code clean, make the tests pass, and if you need to check the current PyTorch or minGPT docs online, go ahead.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab