AutoDAN. [ICLR 2024] The official implementation of our ICLR2024 paper "AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models".
453Universal-Prompt-Injection. The official implementation of our pre-print paper "Automatic and Universal Prompt Injection Attacks against Large Language Models".
73lm-ssp. A reading list for large models safety, security, and privacy.
1LowRankGAN. [NeurIPS 2021] Low-Rank Subspaces in GANs
1pytorch-metric-learning. The easiest way to use deep metric learning in your application. Modular, flexible, and extensible. Written in PyTorch.
1PatchAttack. Python
1evaluation-webpage.
1