nmuru/Executive-Function-Benchmark ? reverse-engineered prompt
Reverse engineered prompt
Build me a simple website and notebook based benchmark that evaluates AI models on executive function using Wordle style tasks and information gain.
I want a clean dashboard that shows a leaderboard, compares models across single turn and multi turn tasks, and highlights things like working memory, cognitive flexibility, and inhibitory control. It should feel interactive, so I can tweak scoring weights and see the rankings and cognitive style update right away. Please also include a way to turn the benchmark outputs into a static site from the evaluation data, since the results come from standardized notebook runs.
Use the existing benchmark data and make the pages easy to read, with sections for the different task types and short explanations of what each metric means. If you need to check any current docs or examples online, go ahead. I mainly want something polished, useful, and ready to publish so people can explore how different models perform on these executive function measures.
Are you gonna build this?
make sure you review the code using coderabbit