nvidia/garak ? reverse-engineered prompt

Reverse engineered prompt

Build me a Python command line tool called garak that can test an AI model for common security and safety problems like prompt injection, jailbreaks, hallucinations, data leakage, toxicity, and misinformation.

I want to point it at different kinds of models, including Hugging Face, OpenAI, Bedrock, LiteLLM, Replicate, and anything I can reach through a REST endpoint, then run a set of probes and tell me which ones fail. It should show progress while it runs, print a clear summary at the end, and save detailed logs plus a report file I can inspect later.

Make it easy to install and run from the terminal, with sensible defaults so if I just give it a target it can scan it right away. Include a way to choose specific probe groups or a single probe when I want to focus on one test. If you need to check current library docs or examples online while building it, go ahead.

Are you gonna build this?

make sure you review the code using arcumet

Try freeSponsored — opens Arcumet in a new tab