Robuta

https://www.harmbench.org/ HarmBench A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal https://openreview.net/forum?id=F0DzK7CIHk GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory | OpenReview Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate... through the lens https://huggingface.co/collections/cais/harmbench-classifiers HarmBench Classifiers - a cais Collection Classifiers for red teaming evaluation in HarmBench classifierscaiscollection