https://arxiv.org/abs/2510.04891v1
[2510.04891v1] SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
Abstract page for arXiv paper 2510.04891v1: SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
revealingllmvulnerabilitiessociallyharmful
https://is.mpg.de/publications/pandeyetal26
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests | MPI-IS
Our goal is to understand the principles of Perception, Action and Learning in autonomous systems that successfully interact with complex environments and to...
revealingllmvulnerabilitiessociallyharmful
https://arxiv.org/abs/2510.04891v2
[2510.04891v2] SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
Abstract page for arXiv paper 2510.04891v2: SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
revealingllmvulnerabilitiessociallyharmful