AI Safety72 / 209 projects
AI safety leaderboard featuring the hottest AI security tool projects on GitHub for red teaming and content filtering.
SafeLine is a self-hosted WAF(Web Application Firewall) / reverse proxy to protect your web apps from attacks and exploits.
🔒 A compiled checklist of 300+ tips for protecting digital security and privacy in 2026
Force Remove Copilot, Recall and More in Windows 11
An open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data. Supports NLP, pattern matching, and customizable pipelines.
This is a multi-use bash script for Linux systems to audit wireless networks.
WiFi security auditing tools suite
NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
Superagent protects your AI applications against prompt injections, data leaks, and harmful outputs. Embed safety directly into your app and prove compliance to your customers.
Passbolt Community Edition (CE) API. The JSON API for the open source password manager for teams!
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
Open-Source Unified Vulnerability Management, DevSecOps & ASPM
A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
CISO Assistant is a one-stop-shop GRC platform for Risk Management, AppSec, Compliance & Audit, TPRM, BIA, Privacy, and Reporting. It supports 150+ global frameworks with automatic control mapping, including ISO 27001, NIST CSF, SOC 2, CIS, PCI DSS, NIS2, DORA, GDPR, HIPAA, CMMC, and more.
scanner detecting the use of JavaScript libraries with known vulnerabilities. Can also generate an SBOM of the libraries it finds.
Open-source security automation platform for teams and AI agents
Sandbox any AI agent in seconds - zero setup, zero latency.
TextAttack 🐙 is a Python framework for adversarial attacks, data augmentation, and model training in NLP https://textattack.readthedocs.io/en/master/
Security scanner for AI agents, MCP servers and agent skills.
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
LLM Prompt Injection Detector
Secrets of RLHF in Large Language Models Part I: PPO
Safety-first guardrails for AI-driven cloud and Kubernetes operations
AI agent security scanner. Detect vulnerabilities in agent configurations, MCP servers, and tool permissions. Available as CLI, GitHub Action, ECC plugin, and GitHub App integration. 🛡️
A security scanner for your LLM agentic workflows
Scan MCP servers for potential threats & security findings.
AI-first security scanner. NEW in v2026.7: Claude Code compromise detection — vet .claude/ hooks, permissions & skills before you clone — plus an always-on AI attack-signature scanner and native Rust & PHP rules. Also: medusa scan --git to vet any repo, medusa secrets scan for leaked API keys. 40,000+ patterns, zero setup.
Open Security Controls Assessment Language (OSCAL)
This is a repository that aims to provide updates on the status of jailbreaking the OpenAI GPT language model.
Open-source AI agent firewall for MCP security and agent egress. Scans mediated HTTP, MCP, A2A, and WebSocket traffic for exfiltration, SSRF, and prompt injection, and emits mediator-signed action receipts: verifiable audit evidence from outside the agent.
CodeGate: Security, Workspaces and Multiplexing for AI Agentic Frameworks
A Deep Graph-based Toolbox for Fraud Detection
A modular, stack-agnostic toolkit of security review skills for AI coding agents to autonomously find, reproduce, and patch vulnerabilities.
AI gets the context. Not your private data. Local-first privacy proxy for browser chat, AI APIs, and coding agents.
Make your GenAI Apps Safe & Secure :rocket: Test & harden your system prompt
High-performance secrets scanner. CLI, Go library, Burp Suite extension, and Chrome extension. 487 detection rules with live credential validation.
VULNRΞPO - Free vulnerability report generator and repository, end-to-end encrypted! Templates of issues, CWE,CVE,MITRE ATT&CK,PCI DSS, import Nmap/Nessus/Burp/OpenVAS/Bugcrowd/Trivy, Jira export, TXT/JSON/MARKDOWN/HTML/DOCX, attachments, automatic changelog, stats, vulnerability management, bugbounty, local ai/llm, super fast pentest reporting!
Deliver safe & effective language models
a browser extension to bring security and privacy to chrome, firefox, and opera
PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. 🏆 Best Paper Awards @ NeurIPS ML Safety Workshop 2022
A robust, and flexible open source User & Entity Behavior Analytics (UEBA) framework used for Security Analytics. Developed with luv by Data Scientists & Security Analysts from the Cyber Security Industry. [BETA]
Security and Privacy Risk Simulator for Machine Learning (arXiv:2312.17667)
World's first Face Authentication enabled MacOS App-locker, completely free and open-source. Unlock your Mac apps using Face , TouchID or password. Completely local and encrypted - your data never leaves your Mac
Simulate a federated setting and run differentially private federated learning.
Backdoors Framework for Deep Learning and Federated Learning. A light-weight tool to conduct your research on backdoors.
Lyrie.ai — The world's first autonomous AI cybersecurity agent. Built by OTT Cybersecurity LLC.
🤖 Admyral enables continuous control monitoring for any custom control
Autonomous white-hat security auditor for AI-driven code review, bug bounty research, exploit construction, and execution-grounded verification.
Breaching privacy in federated learning scenarios for vision and text
Self-hardening firewall for large language models
Agent Execution Partnership AEE is an open-source control plane that ensures every AI agent action is authorized before it runs, observable while it runs, and verifiable after it completes.
BeaverTails is a collection of datasets designed to facilitate research on safety alignment in large language models (LLMs).
A Docker & dev container setup for securely running AI agents in `--dangerous` mode. All container traffic is routed through a transparent mitmproxy, enforcing network access rules and injecting secrets.
Framework for LLM evaluation, guardrails and security
Open-source adversarial testing engine, SDK, and CLI for AI agents. Runs locally or against the Humanbound Platform.
Runtime safety for AI coding agents with real-time enforcement, system-event monitoring, and long-horizon provenance. Supports Claude Code, Codex, Antigravity, Omnigent on native macOS and Linux.
Code Repository for: AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
LLM alignment jailbreak; a set of instructions for auditing their internal reasoning and uncovering biases
AgentGuard: Zero-Trust Security Foundation for AI Agents
Chrome extension risk scanner — scan Chrome Web Store links or CRX/ZIP builds and generate evidence-based security/privacy reports. Open-core.
mcp & skill scanner that scans any mcp server or skills for indirect attack vectors and security or configuration vulnerabilities
A passive way to find backups/ sensitive information.
Repository for "SecurityEval Dataset: Mining Vulnerability Examples to Evaluate Machine Learning-Based Code Generation Techniques" published in MSR4P&S'22.
:zap: Fast Web Security Scanner written in Rust based on Lua Scripts :waning_gibbous_moon: :crab:
A collaborative learning framework for empowering biomedical research
Governance gateway for AI agents — bounded, auditable, session-aware control with MCP proxy, shell proxy & HTTP API. Works with Cursor, Claude Code, Codex, and any MCP-compatible agent.
The open source security engine for AI agent and supply-chain trust.
Statically Detecting Vulnerable Data Flows in Browser Extensions at Scale
🥇 Amazon Nova AI Challenge Winner - ASTRA emerged victorious as the top attacking team in Amazon's global AI safety competition, defeating elite defending teams from universities worldwide in live adversarial evaluation.
RAG/LLM Security Scanner identifies critical vulnerabilities in AI-powered applications, including chatbots, virtual assistants, and knowledge retrieval systems.
Automatic security vulnerability remediation for your code.
Autonomous Bug Bounty Hunting Framework powered by Claude Code. 20 AI agents, state-machine orchestration, Burp Suite MCP, credential vault, LLM security track. Type 'hunt target.com' and let AI find the bugs.
A cutting-edge, high-tech autonomous banking network built with advanced security features, robust authentication, and authorization mechanisms. The GBN repository includes a modular architecture, secure communication protocols, and a robust incident response plan, making it a reliable and secure platform for financial transactions in the galaxy.
Source: GitHub API · Curated · Realtime