UKGovernmentBEIS/inspect_ai ? reverse-engineered prompt
Reverse engineered prompt
Build me a Python framework for evaluating large language models.
I want a tool that lets me define evals, run prompts through models, support tool use and multi turn conversations, and score the results automatically, including model graded judging. It should come with a bunch of ready to run example evaluations so I can test different models right away, and it should be easy to add my own custom evals, scoring rules, and extensions later.
Please make it feel like a real developer friendly library, with clear docs, a simple command line workflow, and a small web UI if that fits naturally. I’d also like the project set up so it’s easy to install in editable mode for local development, with tests, linting, formatting, and a clean path for contributors to work on it. If you need to check current docs or best practices while building, look them up online.
Are you gonna build this?
make sure you review the code using coderabbit