https://openreview.net/forum?id=ZaOHSBGOhV&referrer=%5Bthe%20profile%20of%20Jiaxin%20Wen%5D(%2Fprofile%3Fid%3D~Jiaxin_Wen2)
SmartBackdoor: Malicious Language Model Agents that Avoid Being Caught | OpenReview
As large language model (LLM) agents receive more information about themselves or the users from the environment, we speculate a new family of cyber attacks,...
language modelbeing caughtmaliciousagentsavoid