qatration/qatration ? reverse-engineered prompt

Reverse engineered prompt

Build me a Python tool that I can run on my own machine or in CI to red team my own chatbot or agent deployment.

I want to point it at a chat endpoint, have it send a library of real attacks like prompt injection, system prompt leakage, data exfiltration, tool abuse, and agent style SQL injection, then tell me clearly which ones worked and which ones were defended. Please make the results objective, using planted canaries and tool call checks instead of a model judge, and save evidence locally in a results folder that can be turned into SARIF for CI.

It should have a simple setup flow that writes a target config for me, mints a canary, and lets me onboard the target before running the full sweep. It should also support a benign mode for seeing what happens with normal traffic, and it should refuse to run if the config is missing the canary or if the target is not authorised. If you need to look up current docs online, go ahead.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab