harveyai/harvey-labs ? reverse-engineered prompt

Reverse engineered prompt

Build me an open source benchmark for testing AI agents on real legal work.

I want a project that includes a set of legal tasks with instructions, source documents, and clear scoring rubrics, plus a way to run an agent on those tasks and evaluate the results. It should support realistic workflows like reviewing documents, following instructions in a data room style assignment, and comparing different agent runs so I can see which one performs better.

Please include a simple walkthrough so someone can set it up, inspect a task, run an agent, score the output, and review the report. Make it easy to add new tasks and improve the evaluation over time. If you need to check current best practices or docs while building it, go ahead and look them up online.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab