This is your work, valued
I-GCG. Improved techniques for optimization-based jailbreaking on large language models (ICLR2025)
★ 146LAS-AT. Code for LAS-AT: Adversarial Training with Learnable Attack Strategy (CVPR2022)
★ 120Comdefend. The code for ComDefend: An Efficient Image Compression Model to Defend Adversarial Examples (CVPR2019)
★ 114OmniSafeBench-MM. A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack–Defense Evaluation
★ 75SkillJect. SkillJect: Automating Stealthy Skill-Based Prompt Injection for Coding Agents with Trace-Driven Closed-Loop Refinement
★ 73FOA-Attack. Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment (NeurIPS 2025)
★ 67FGSM-PGK. Improving fast adversarial training with prior-guided knowledge (TPAMI2024)
★ 43SA-AET. Code for Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language Attack(TPAMI 2025)
★ 42Adv-watermark. Code for Adv-watermark: A novel watermark perturbation for adversarial examples (ACM MM2020)
★ 40FGSM-LAW. Revisiting and Exploring Efficient Fast Adversarial Training via LAW: Lipschitz Regularization and Auto Weight Averaging (TIFS2024)
★ 37FGSM-SDI. Code for Boosting fast adversarial training with learnable adversarial initialization (TIP2022)
★ 29FGSM-PGI. Code for Prior-Guided Adversarial Initialization for Fast Adversarial Training (ECCV2022)
★ 28FP-Better. Code for Fast Propagation is Better: Accelerating Single-Step Adversarial Training via Sampling Subnetworks (TIFS2024)
★ 13Robust_Logo_Detection. object detection; robust detection; ACM MM21 grand challenge; Security AI Challenger Phase VII
★ 3jiaxiaojunQAQ.github.io. AcadHomepage: A Modern and Responsive Academic Personal Homepage
★ 1Watermark-Vaccine. The code for ECCV2022 (Watermark Vaccine: Adversarial Attacks to Prevent Watermark Removal)
★ 1ER-APT.
★ 1llm-fingerprinting. single-token behavioural fingerprints (arXiv:2607.10252) One Token Is Enough
★ 24Awesome-Self-Evolving-Agents. A Survey of Self-Evolving Agents | A curated list of resources (surveys, papers, benchmarks, and opensource projects) on Self-Evolving Agents.
★ 360personal-model. Build your HUMAN.md.
★ 1.3kNoise-Robust-Density-Estimation. Python
★ 1Audar-ASR-V1. Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
★ 560BettaFish. 微舆:人人可用的多Agent舆情分析助手,打破信息茧房,还原舆情原貌,预测未来走向,辅助决策!从0实现,不依赖任何框架。
★ 42kAP2. Building a Secure and Interoperable Future for AI-Driven Payments.
★ 3.1kDeepScan. Official implement for "DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models"
★ 242gmixer. [CVPR 2026] G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
★ 8Chinese-offensive-language-detect. 本项目旨在通过数据生成技术、模型微调技术、加权投票技术来检测以下6种有害文本:涉黄有害文本、辱骂/谐音辱骂有害文本、地域、性别、种族、职业攻击有害文本
★ 868Arbor. A generalist autonomous research agent — runs experiments, researches, and iteratively optimizes, autonomously.
★ 978CCFA-Skills. A skill family for shaping the research storyline of CCF-A papers.
★ 1kRECODE-H. Python
★ 6HAMLET-Isaac-GR00T. HAMLET + Isaac GR00T (N1.6/N1.5)
★ 29RobustVLA. Python
★ 21RodriguesNetwork. [ICLR 2026 Oral] Rodrigues Network for Learning Robot Actions
★ 25secureclaw. SecureClaw - Security Plugin and Skill for OpenClaw OWASP-Aligned
★ 349Agent3Sigma-Canary. Agent3σ-Canary is an evaluation framework for AI Agent security in realistic runtime environments.
★ 33ACoT-VLA. [CVPR 2026] Official implementation of "ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models"
★ 224starVLA. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
★ 3.3kUniRL. UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
★ 863Anomaly-OneVision. This is the official repository for our recent paper "Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models".
★ 171INP-Former. [CVPR 2025] official implementation of “Exploring Intrinsic Normal Prototypes within a Single Image for Universal Anomaly Detection”
★ 308hazardarena-anon. HazardArena: Evaluating Semantic Safety in VLA Models
★ 2SafeVLA. [NeurIPS 2025 Spotlight] Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning.
★ 154agentic_geo. Code for AgenticGEO: A Self-Evolving Agentic System for Generative Engine Optimization
★ 36AE-CoT. Python
★ 3AgentDoG. A Diagnostic Guardrail Framework for AI Agent Safety and Security
★ 671Sensitive-lexicon. 一个持续更新的中文敏感词库,帮助开发者和内容审核者快速识别并过滤不当文本,即将迎来重大更新
★ 3.9kAutoFigure-Edit. Python
★ 4kMMBench-GUI. Official repo of "MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents". It can be used to evaluate a GUI agent with a hierarchical manner across multiple platforms, including Windows, Linux, macOS, iOS, Android and Web.
★ 113mlrm-LEAD. [CVPR 2026 Highlight] Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware Decoding
★ 95Awesome-Auto-Research-Tools. A curated collection of automated research tools, covering literature search, paper reading, experiment management, and code generation to help researchers accelerate their workflow.
★ 1.1kTool-Star. 🔧Tool-Star: Empowering LLM-brained Multi-Tool Reasoner via Reinforcement Learning
★ 405ToolRL. Python
★ 514C2P-CLIP-DeepfakeDetection. C2P-CLIP-DeepfakeDetection
★ 102Auto-claude-code-research-in-sleep. ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
★ 14kMPCAttack. [CVPR 2026] Multi-Paradigm Collaborative Adversarial Attack Against Multimodal Large Language Models
★ 12DARWIN. DARWIN is a self-evolving LLM jailbreak framework that grows a reusable strategy pool through external extraction, sandbox filtering, history-memory retrieval, Markov-based strategy selection, and reflective/genetic evolution.
★ 61GLPN-LLM. ACL 2025 Main: synergizing llms with global label propagation for multimodal fake news detection
★ 41AutoSOTA. Python
★ 617claude-mem. Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
★ 89kdarwin-skill. 达尔文.skill —— 一个让你的Skill无限进化的系统:评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.
★ 5.2kautoresearch. AI agents running research on single-GPU nanochat training automatically
★ 92kOpenHarness. "OpenHarness: Open Agent Harness with a Built-in Personal Agent--Ohmo!"
★ 15kClaude-Code-Source-Study. Deep dive into Claude Code's source code— learn from the best agent implementation out there.
★ 1.6kex-skill. 致你忘不掉的那个TA,你们干大模型都是码圣 It's giving rebirth era. Welcome to Digital Life 1.0. 🫶
★ 1kagentic-harness-patterns-skill. Agent skill for harness engineering — memory, permissions, context engineering, multi-agent coordination. Distilled from Claude Code, with Codex CLI and Gemini CLI on the roadmap. EN/ZH. Install via npx skills add.
★ 299claw-code. An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
★ 195kAwesome-Embodied-AI-Safety. Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses | 500+ Papers | Perception, Cognition, Planning, Interaction, Agentic System
★ 120ClawKeeper. ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers (aka The Norton for OpenClaw)
★ 1kslowmist-agent-security. SlowMist Agent Security Skill: A comprehensive security review framework for AI agents operating in adversarial environments. Core principle: Every external input is untrusted until verified.
★ 495pypi_malregistry. The repository has collected over 10,000 malicious pypi packages. This dataset is the work of the ASE 2023 paper "An Empirical Study of Malicious Code In PyPI Ecosystem". Of course, we will continue to expand the dataset. Latest update time: 21 July. 2026
★ 128PRISM. Codes for paper Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
★ 8inception-T2I-system.
★ 7clawguard. JavaScript
★ 50MaliciousAgentSkillsBench. A Security Benchmark for Claude Code Agent Skills
★ 69agent-security-skill-scanner. Python
★ 26awesome-red-teaming-llms. Papers from our SoK on Red-Teaming (Accepted at TMLR)
★ 45agent-scan. Security scanner for AI agents, MCP servers and agent skills.
★ 2.8kskill-inject. Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks
★ 88CLI-Anything. "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
★ 46kAegis. Runtime policy enforcement for AI agents. Cryptographic audit trail, human-in-the-loop approvals, kill switch. Zero code changes.
★ 371CC-BOS. Python
★ 244openclaw-eval. Python
★ 27locomo. Python
★ 1.1kOpenViking. Self-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.
★ 28kagent-audit. Static security scanner for LLM agents — prompt injection, MCP config auditing, taint analysis. 51 rules mapped to OWASP Agentic Top 10 (2026). Works with LangChain, CrewAI, AutoGen.
★ 205Agent-Reach. Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
★ 62kclawdbot_report.
★ 12ClawX. ClawX is a desktop app that provides a graphical interface for OpenClaw AI agents. It turns CLI-based AI orchestration into a desktop experience without using the terminal. China website is https://clawx.com.cn.
★ 7.6kskill-scanner. Security Scanner for Agent Skills
★ 2.4kLLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9kToolathlon. [ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
★ 443awesome-openclaw-skills. The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
★ 52ktrustclaw. TypeScript
★ 49marketplace. Security-audited skills for Claude, Codex & Claude Code. One-click install, quality verified.
★ 405openclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 384kwooyun-legacy. wooyun-legacy skill for claude code
★ 1.7kBackdoorAgent. BackdoorAgent is a stage-aware framework and benchmark that instruments LLM-agent workflows (planning, memory, tools) to systematically inject, track, and evaluate backdoor triggers across multi-step trajectories using unified metrics like ASR and clean accuracy.
★ 43opencode-skillful. OpenCode Skills Plugin - Allow your opencode agents to lazy load prompts on demand.
★ 317BingoGuard. Python
★ 15OmniSafeBench-MM. A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack–Defense Evaluation
★ 75skills. Public repository for Agent Skills
★ 165kUltraBr3aks. sharing NEW strong AI jailbreaks of multiple vendors (LLMs)
★ 370personality_in_llms. Jupyter Notebook
★ 49propensity-evaluation. open Source code for propensity evaluation
★ 19UniCardio. Versatile Cardiovascular Signal Generation with a Unified Diffusion Transformer
★ 62Odysseus. [NDSS 2026] Official repo for Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography
★ 59EEG-To-Text. EEG-To-Text
★ 4Thought2Text. EEG to Language through LLMs
★ 122SPELL. This is the code for the FSE26 paper: Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking, the research paper can be referred at https://arxiv.org/pdf/2512.21236
★ 4physpatch. The official implementation of PhysPatch; this paper was accepted by AAAI 2026.
★ 9LLMs-Zero-to-Hero. 从无名小卒到大模型(LLM)大英雄~ 欢迎关注后续!!!
★ 2.3kGeoshield. Official PyTorch implementation of AAAI 2026 paper Geoshield
★ 13ToxiCN_MM. The code and resource of "Towards Comprehensive Detection of Chinese Harmful Memes" (NeurIPS2024 D&B).
★ 86speechBCI. Code for the paper "A high-performance speech neuroprosthesis"
★ 209agents. An Open-source Framework for Data-centric, Self-evolving Autonomous Language Agents
★ 6kPLeak. Python
★ 82AgentSafety.
★ 192MCIP. Python
★ 13code. Python
★ 48openinterpreter. A coding agent for open models like Kimi K3
★ 67kMCP-Security-Checklist. A comprehensive security checklist for MCP-based AI tools. Built by SlowMist to safeguard LLM plugin ecosystems.
★ 835CoPS. [CVPR 2026 Findings] Official Implementation for "CoPS: Conditional Prompt Synthesis for Zero-Shot Anomaly Detection"
★ 50LogSAD. [CVPR 2025] Towards Training-free Anomaly Detection with Vision and Language Foundation Models
★ 101Misevolution. Official Repo of Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
★ 90PBI-Attack. Python
★ 6mstar. [ICML 2025] M-STAR (Multimodal Self-Evolving TrAining for Reasoning) Project. Diving into Self-Evolving Training for Multimodal Reasoning
★ 75rStar. Python
★ 1.4kHiddenDetect. ACL 2025 (Main) HiddenDetect: Detecting Jailbreak Attacks against Multimodal Large Language Models via Monitoring Hidden States
★ 165CS-DJ. Accept by CVPR 2025 (highlight)
★ 25Oyster. The Oyster series is a set of safety models developed in-house by Alibaba-AAIG, devoted to building a responsible AI ecosystem. | Oyster 系列是 Alibaba-AAIG 自研的安全模型,致力于构建负责任的 AI 生态。
★ 62DH-CoT. Two text jailbreak attacks against commercial black-box LLMs and a malicious content detection method, the latter of which is applied to red team dataset cleaning and jailbreak response detection.
★ 12awesome-llm-apps. 100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
★ 129kStrata-Sword. The Strata-Sword is a hierarchical Chinese-English jailbreak safety benchmark based on quantified reasoning complexity, developed in-house by Alibaba-AAIG | Strata-Sword 是 Alibaba-AAIG自研的中英文分层越狱攻击安全基准,将“推理复杂度”作为可评估的安全维度,并提出多种中文特有攻击方法,以系统评测不同推理复杂度下LLMs和LRMs的安全边界,从而为提升模型安全性提供新思路。
★ 22python-sdk. The official Python SDK for Model Context Protocol servers and clients
★ 24kuniversal-mcp. Universal MCP acts as a middle ware for your API applications. It can store your credentials, authorize, enable disable apps on the fly and much more.
★ 26AI-Infra-Guard. A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
★ 4.3kAutoRAN-public. Python
★ 11DeepResearch. Tongyi Deep Research, the Leading Open-source Deep Research Agent
★ 20kOpenSafety. Open Safety Gym with PyBullet
★ 8DrAttack. Official implementation of paper: DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
★ 68OpenAgentSafety. A Framework for Evaluating AI Agent Safety in Realistic Environments
★ 38red-teaming-challenge. Red‑Teaming Challenge - OpenAI gpt-oss-20b
★ 7gpt-oss. gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
★ 20kSafeAgentBench. Codes for paper "SafeAgentBench: A Benchmark for Safe Task Planning of \\ Embodied LLM Agents"
★ 74CognitiveOverload. Code for our NAACL 2024 Paper "Cognitive Overload: Jailbreaking Large Language Models with Overloaded Logical Thinking"
★ 8learn-claude-code. Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
★ 73kJOOD. [CVPR 2025] Official implementation for JOOD "Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy"
★ 21PromptInject. PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. 🏆 Best Paper Awards @ NeurIPS ML Safety Workshop 2022
★ 513R-2-Guard. [ICLR 2025] Code implementation of R^2-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical Reasoning
★ 23Agent-S. Agent S: an open agentic framework that uses computers like a human
★ 12kZetaLib. 🌙 ZetaLib - The only AI Library you need
★ 826Algo-Trade-Adversarial-Examples. todo: desc
★ 11EmojiAttack. Emoji Attack [ICML 2025]
★ 46AsFT. Code for the paper "AsFT: Anchoring Safety During LLM Fune-Tuning Within Narrow Safety Basin".
★ 37UDora. [ICML 2025] UDora: A Unified Red Teaming Framework against LLM Agents
★ 38agentdojo. A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.
★ 694STAIR. Official codebase for "STAIR: Improving Safety Alignment with Introspective Reasoning"
★ 89ArtPrompt. [ACL24] Official Repo of Paper `ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs`
★ 102CodeAttack. [ACL 2024] CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
★ 62panda-guard. Panda Guard is designed for researching jailbreak attacks, defenses, and evaluation algorithms for large language models (LLMs).
★ 69awesome-mcp-servers. A collection of MCP servers.
★ 92kFlipAttack. [ICML 2025] An official source code for paper "FlipAttack: Jailbreak LLMs via Flipping".
★ 179mcpSafetyScanner. MCPSafetyScanner - Automated MCP safety auditing and remediation using Agents. More info: https://www.arxiv.org/abs/2504.03767
★ 176LLMLingua. [EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
★ 6.5kDeep-RL-for-Automated-Stock-Trading. Jupyter Notebook
★ 13advertorch. A Toolbox for Adversarial Robustness Research
★ 1.4kX-Guard-Multilingual-Guard-Agent-for-Content-Moderation. Jupyter Notebook
★ 2NBF-LLM. The official code for "Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks".
★ 18Reasoning-to-Defend. [EMNLP 2025] Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking
★ 12ICRT. Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
★ 4AwesomeLLMJailBreakPapers. Awesome LLM Jailbreak academic papers
★ 168SurgVLM. CSS
★ 66PocketGen. PocketGen (Nature Machine Intelligence 24): Generating Full-Atom Ligand-Binding Protein Pockets
★ 224x-teaming. Python
★ 67MTSA. offical implementation of MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming
★ 17HalluLens. Codebase for LLM Textual Hallucination Benchmark
★ 84XTransferBench. [ICML 2025] X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
★ 48UnivIntruder. Python
★ 17Awesome-LVLM-Attack. 😎 up-to-date & curated list of awesome Attacks on Large-Vision-Language-Models papers, methods & resources.
★ 566AudioTrust. AudioTrust: Benchmarking the Multi-faceted Trustworthiness of Audio Large Language Models
★ 216Stochastic-Gradient-Aggregation. Official implementation of the ICCV2023 paper: Enhancing Generalization of Universal Adversarial Perturbation through Gradient Aggregation
★ 28DM-UAP. [AAAI 2025] Improving Generalization of Universal Adversarial Perturbation via Dynamic Maximin Optimization
★ 6Navig. Python
★ 16Efficient-Vision-Language-Models-A-Survey. [2025] Efficient Vision Language Models: A Survey
★ 52PackageHallucination. Code and data for the USENIX 2025 paper "We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs"
★ 32TransferAttack. [ACL 2025] Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
★ 19Adversarial-Reasoning. A new algorithm that formulates jailbreaking as a reasoning problem.
★ 26SafeRAG. Python
★ 62VLSBench. [ACL 2025] Data and Code for Paper VLSBench: Unveiling Visual Leakage in Multimodal Safety
★ 62privacy-inference-multimodal. Python
★ 21ActorAttack. Python
★ 134Awesome-Multi-Turn-LLMs. This is the official GitHub repository for our survey paper "Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models".
★ 201Foot-in-the-door-Jailbreak. Python
★ 23langchain. The agent engineering platform.
★ 143kawesome-foundation-agents. About Awesome things towards foundation agents. Papers / Repos / Blogs / ...
★ 2.2kAutomated-Multi-Turn-Jailbreaks. Python
★ 139multi-turn-attack-defenses. Python
★ 11jailbreak-reasoning-openai-o1o3-deepseek-r1. Jupyter Notebook
★ 121Embodied-AI-Guide. [Lumina具身智能社区] 具身智能技术指南 Embodied-AI-Guide
★ 15kunthinking_vulnerability. To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
★ 33awesome-computer-use. This is a collection of resources for computer-use GUI agents, including videos, blogs, papers, and projects.
★ 572factif-ai. AI-powered computer control for automated testing. Factifai uses vision models (Claude, GPT-4o, Gemini) to interact with applications naturally - clicking, typing, and verifying results just like a human would.
★ 56MMVP. Python
★ 365vlm-evaluation. VLM Evaluation: Benchmark for VLMs, spanning text generation tasks from VQA to Captioning
★ 139Robust-LLaVA. [ICCVW 2025 (Oral)] Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
★ 29MOSSBench. An implementation for MLLM oversensitivity evaluation
★ 18ZSRobust4FoundationModel. Python
★ 48M-Attack. [NeurIPS25 & ICML25 Workshop on Reliable and Responsible Foundation Models] A Simple Baseline Achieving Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1. Paper at: https://arxiv.org/abs/2503.10635
★ 100HySAC. Hyperbolic Safety-Aware Vision-Language Models. CVPR 2025
★ 31Awesome-Multimodal-Jailbreak. A Survey on Jailbreak Attacks and Defenses against Multimodal Generative Models
★ 333perceptual-metrics. Python
★ 8tg_gbc. [ICCV2025] Accelerate 3D Object Detection Models via Zero-Shot Attention Key Pruning
★ 11AI-Supervision-Risk.
★ 21magic. The repo for paper: Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models.
★ 15SIAGT. Python
★ 4Immune. [CVPR2025] Official Repository for IMMUNE: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
★ 28Awesome-Multimodal-Large-Language-Models. 🔥Awesome Multimodal Large Language Models Paper List
★ 154