uiuc-kang-lab/cve-bench ? reverse-engineered prompt
Reverse engineered prompt
Build me a Python tool that runs a benchmark for testing whether AI agents can exploit real web app vulnerabilities. It should let me pull the challenge containers, start and stop a task, run an evaluation against a model, retry failed runs, and generate the default prompt for each challenge. I also want a way to test the health of the tasks and run the provided solution version so I can check that everything works end to end.
Use Docker for the apps and keep the setup reproducible. Make it easy to run from one main command, and include support for selecting specific challenges and variants instead of always running everything. The prompts and metadata for each CVE should come from simple files so it’s easy to add more later. If you need current details for the evaluation framework or Docker commands, look them up online.
Are you gonna build this?
make sure you review the code using coderabbit