Roblox/RobloxGuard-1.0 ? reverse-engineered prompt

Reverse engineered prompt

Build me a simple Python tool that uses Roblox Guard 1.0 to check whether a user prompt or a generated response is unsafe.

I want to be able to point it at a config file, load the matching prompt template, run inference with the Roblox Guard model, and save the results to CSV. It should support evaluating a dataset from Hugging Face, compare the model prediction against the ground truth label when one is available, and also write a summary file with counts and basic metrics like precision, recall, F1, and false positive rate.

Please make it easy to swap in a different taxonomy or dataset by editing the config and prompt text, and keep the output format clear so I can inspect each example. If anything is unclear, look up the current model and dataset docs online and make the script work with them.

Are you gonna build this?

make sure you review the code using coderabbit

Try freeSponsored — opens CodeRabbit in a new tab