Security & Governance413 projects
A curated leaderboard of AI security and governance projects on GitHub, including AI safety, model evaluation, privacy, and explainability tools.
This repository is maintained by Omar Santos (@santosomar) and includes thousands of resources related to ethical hacking, bug bounties, digital forensics and incident response (DFIR), AI security, vulnerability research, exploit development, reverse engineering, and more. 🔥 Also check: https://hackertraining.org
Fully automatic censorship removal for language models
A game theoretic approach to explain the output of any machine learning model.
FHEVM, a full-stack framework for integrating Fully Homomorphic Encryption (FHE) with blockchain applications
SafeLine is a self-hosted WAF(Web Application Firewall) / reverse proxy to protect your web apps from attacks and exploits.
Identity-aware VPN and tunneled reverse proxy for remote access based on WireGuard®.
🔒 A compiled checklist of 300+ tips for protecting digital security and privacy in 2026
The LLM Evaluation Framework
Python SDK for Agent AI Observability, Monitoring and Evaluation Framework. Includes features like agent, llm and tools tracing, debugging multi-agentic system, self-hosted dashboard and advanced analytics with timeline and execution graph view
Zama Bounty Program: Contribute to the FHE space and Zama's open source libraries and get rewarded 💰
Supercharge Your LLM Application Evaluations 🚀
Extract and decrypt browser data, supporting multiple data types, runnable on various operating systems (macOS, Windows, Linux).
Security scanner for AI agent skills. Detect vulnerabilities, malicious patterns, security risks, prompt injection, data exfiltration, and supply-chain risks in Claude Code, Codex, and MCP skills before you install them.
A framework for few-shot evaluation of language models.
Bitwarden client apps (web, browser extension, desktop, and cli).
Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
Force Remove Copilot, Recall and More in Windows 11
Lime: Explaining the predictions of any machine learning classifier
AI Observability & Evaluation
An open-source framework for detecting, redacting, masking, and anonymizing sensitive data (PII) across text, images, and structured data. Supports NLP, pattern matching, and customizable pipelines.
Perform data science on data that remains in someone else's server
Cybersecurity AI (CAI), the framework for AI Security
Transformer Explained Visually: Learn How LLM Transformer Models Work with Interactive Visualization
BertViz: Visualize Attention in Transformer Models
OpenShell is the safe, private runtime for autonomous AI agents.
This is a multi-use bash script for Linux systems to audit wireless networks.
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
🛡️ Windows Hello™ style facial authentication for Linux
😈Awful AI is a curated list to track current scary usages of AI - hoping to raise awareness
WiFi security auditing tools suite
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Adding guardrails to large language models.
Fit interpretable models. Explain blackbox machine learning.
NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
Superagent protects your AI applications against prompt injections, data leaks, and harmful outputs. Embed safety directly into your app and prove compliance to your customers.
Adversarial Robustness Toolbox (ART) - Python Library for Machine Learning Security - Evasion, Poisoning, Extraction, Inference - Red and Blue Teams
An Industrial Grade Federated Learning Framework
Passbolt Community Edition (CE) API. The JSON API for the open source password manager for teams!
Democratizing Reinforcement Learning for LLMs
AI Agent Governance Toolkit — Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10 OWASP Agentic Top 10.
A huge blocklist of manually curated sites that contain AI generated imagery for uBlock Origin & uBlacklist.
Robust recipes to align language models with human and AI preferences
SWE-bench: Can Language Models Resolve Real-world Github Issues?
Fawkes, privacy preserving tool against facial recognition systems. More info at https://sandlab.cs.uchicago.edu/fawkes
autonomous red teaming platform; multi-agent offensive-security meta-harness
Adversary simulation and Red teaming platform with AI
Stalk your Friends. Find their Instagram, FB and Twitter Profiles using Image Recognition and Reverse Image Search.
Autonomous Hacking Agent for Red Team
Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud, no patient data leaving your network. Apache-2.0
《可解释的机器学习--黑盒模型可解释性理解指南》,该书为《Interpretable Machine Learning》中文版
Open-Source Unified Vulnerability Management, DevSecOps & ASPM
Universal and Transferable Attacks on Aligned Language Models
A collection of infrastructure and tools for research in neural network interpretability.
Align Anything: Training All-modality Model with Feedback
A full-stack AI Red Teaming platform securing AI ecosystems via OpenClaw Security Scan, Agent Scan, Skills Scan, MCP scan, AI Infra scan and LLM jailbreak evaluation.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
CISO Assistant is a one-stop-shop GRC platform for Risk Management, AppSec, Compliance & Audit, TPRM, BIA, Privacy, and Reporting. It supports 150+ global frameworks with automatic control mapping, including ISO 27001, NIST CSF, SOC 2, CIS, PCI DSS, NIS2, DORA, GDPR, HIPAA, CMMC, and more.
中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard
The Python Risk Identification Tool for generative AI (PyRIT) is an open source framework built to empower security professionals and engineers to proactively identify risks in generative AI systems.
scanner detecting the use of JavaScript libraries with known vulnerabilities. Can also generate an SBOM of the libraries it finds.
Open Source Data Security Platform for Developers to Monitor and Detect PII, Anonymize Production Data and Sync it across environments.
A Web app to manage your Two-Factor Authentication (2FA) accounts and generate their security codes
Deepchecks: Tests for Continuous Validation of ML Models & Data. Deepchecks is a holistic open-source solution for all of your AI & ML validation needs, enabling to thoroughly test your data and models from research to production.
A list of AI agents and robots to block.
[CCS'24] A dataset consists of 15,140 ChatGPT prompts from Reddit, Discord, websites, and open-source datasets (including 1,405 jailbreak prompts).
Open-source security automation platform for teams and AI agents
A library for mechanistic interpretability of GPT-style language models
Automation for internal Windows Penetrationtest / AD-Security
The Learning Interpretability Tool: Interactively analyze ML models to understand their behavior in an extensible and framework agnostic interface.
Sandbox any AI agent in seconds - zero setup, zero latency.
OML 1.0 via Fingerprinting: Open, Monetizable, and Loyal AI
Evaluation and Tracking for LLM Experiments and AI Agents
TextAttack 🐙 is a Python framework for adversarial attacks, data augmentation, and model training in NLP https://textattack.readthedocs.io/en/master/
Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents
Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
The Security Toolkit for LLM Interactions
Privacy-first password manager with built-in email aliasing. Fully encrypted and self-hostable.
Neural network visualization toolkit for keras
A Python toolbox to create adversarial examples that fool neural networks in PyTorch, TensorFlow, and JAX
[ACL 2024] An Easy-to-use Knowledge Editing Framework for LLMs.
Security scanner for AI agents, MCP servers and agent skills.
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducible and transparent evaluation of foundation models, including large language models (LLMs) and multimodal models.
A unified evaluation framework for large language models
A library for debugging/inspecting machine learning classifiers and explaining their predictions
A unified framework for privacy-preserving data analysis and machine learning
Algorithms for outlier, adversarial and drift detection
Inspect: A framework for large language model evaluations
UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.
Explain, analyze, and visualize NLP language models. Ecco creates interactive visualizations directly in Jupyter notebooks explaining the behavior of Transformer-based language models (like GPT2, BERT, RoBERTA, T5, and T0).
Source code about machine learning and security.
Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪
Human preference data for "Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback"
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Information Gathering Instagram.
Convert Apple NeuralHash model for CSAM Detection to ONNX.
LLM Prompt Injection Detector
Secrets of RLHF in Large Language Models Part I: PPO
Human ChatGPT Comparison Corpus (HC3), Detectors, and more! 🔥
Model explainability that works seamlessly with 🤗 transformers. Explain your transformers model in just 2 lines of code.
Privacy-Preserving Computing Platform 由密码学专家团队打造的开源隐私计算平台,支持多方安全计算、联邦学习、隐私求交、匿踪查询等。
OWASP Top 10 for Large Language Model Apps (Part of the GenAI Security Project)
Safety-first guardrails for AI-driven cloud and Kubernetes operations
DEPRECATED - Client for Privacy Pass protocol providing unlinkable cryptographic tokens
a security scanner for custom LLM applications
AttackGen is a cybersecurity incident response testing tool that leverages the power of large language models and the comprehensive MITRE ATT&CK framework. The tool generates tailored incident response scenarios based on user-selected threat actor groups and your organisation's details.
Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。
The popular NoScript Security Suite browser extension.
Evaluate your LLM's response with Prometheus and GPT4 💯
An AI-powered threat modeling tool that leverages OpenAI's GPT models to generate threat models for a given application based on the STRIDE methodology.
Benchmarking Generalized Out-of-Distribution Detection
A framework for the evaluation of autoregressive code generation language models.
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Reinforcement Learning via Self-Distillation (SDPO)
Data Selfie - a browser extension to track yourself on Facebook and analyze your data.
AI agent security scanner. Detect vulnerabilities in agent configurations, MCP servers, and tool permissions. Available as CLI, GitHub Action, ECC plugin, and GitHub App integration. 🛡️
A security scanner for your LLM agentic workflows
The nnsight package enables interpreting and manipulating the internals of deep learned models.
Scan MCP servers for potential threats & security findings.
Framework-agnostic implementation for state-of-the-art saliency methods (XRAI, BlurIG, SmoothGrad, and more).
:art: A CNN visualizer
Agent-native reverse-engineering lab with a 197-article knowledge base, MCP tools, and CTF/APK/PE automation workflows.
🔍 LangKit: An open-source toolkit for monitoring Large Language Models (LLMs). 📚 Extracts signals from prompts & responses, ensuring safety & security. 🛡️ Features include text quality, relevance metrics, & sentiment analysis. 📊 A comprehensive tool for LLM observability. 👀
Offen Fair Web Analytics
OmniXAI: A Library for eXplainable AI
Password-protect URLs using AES in the browser; create hidden bookmarks without a browser extension
AI-Native Risk Intelligence Systems, OpenDeRisk——Your application system risk intelligent manager provides 7* 24-hour comprehensive and in-depth protection.
AI-first security scanner. NEW in v2026.7: Claude Code compromise detection — vet .claude/ hooks, permissions & skills before you clone — plus an always-on AI attack-signature scanner and native Rust & PHP rules. Also: medusa scan --git to vet any repo, medusa secrets scan for leaked API keys. 40,000+ patterns, zero setup.
Open Security Controls Assessment Language (OSCAL)
FacTool: Factuality Detection in Generative AI
This is a repository that aims to provide updates on the status of jailbreaking the OpenAI GPT language model.
Official repository for the paper "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code"
Safety guardrails for ai coding agents and human terminal commands
Winner of Mozilla's $50,000 prize for AI
Diffprivlib: The IBM Differential Privacy Library
[ICCV 2021- Oral] Official PyTorch implementation for Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder Transformers, a novel method to visualize any Transformer-based network. Including examples for DETR, VQA.
A deliberately vulnerable banking application designed for practicing Security Testing of Web App, APIs, AI integrated App and secure code reviews. Features common vulnerabilities found in real-world applications, making it an ideal platform for security professionals, developers, and enthusiasts to learn pentesting and secure coding practices.
An easy-to-use Python framework to generate adversarial jailbreak prompts.
Repository for the paper "Automated Hate Speech Detection and the Problem of Offensive Language", ICWSM 2017
A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.
A filter list that hides website features which use Generative AI & content labeled as AI Generated.
Metis is an open-source, AI-driven tool for deep security code review
Autonomous AI pentesting engine, continuous offensive security across web, cloud, identity, CI/CD, IaC, databases, Active Directory, Kubernetes and IoT firmware. Agentic reasoning plus real exploit execution deliver proof-based vulnerabilities. Privacy gateway: the LLM never sees your real IPs, hosts or creds, nothing leaves your perimeter.
The repository is for safe reinforcement learning baselines.
AIRecon is an autonomous cybersecurity agent that combines a self-hosted Large Language Model (Ollama) with a Kali Linux Docker sandbox and a Textual TUI. It is designed to automate security assessments, penetration testing, and bug bounty reconnaissance — without any API keys or cloud dependency.
Disconnect is a browser extension that makes the web faster, more private, and more secure.
Diffusion attentive attribution maps for interpreting Stable Diffusion.
一个浏览器数据(密码|历史记录|Cookie|书签|下载记录)的导出工具,支持主流浏览器。
BLEURT is a metric for Natural Language Generation based on transfer learning.
Open-source AI agent firewall for MCP security and agent egress. Scans mediated HTTP, MCP, A2A, and WebSocket traffic for exfiltration, SSRF, and prompt injection, and emits mediator-signed action receipts: verifiable audit evidence from outside the agent.
CodeGate: Security, Workspaces and Multiplexing for AI Agentic Frameworks
An Open-Source Package for Textual Adversarial Attack.
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
Locating and editing factual associations in GPT (NeurIPS 2022)
BruteSploit is a collection of method for automated Generate, Bruteforce and Manipulation wordlist with interactive shell. That can be used during a penetration test to enumerate and maybe can be used in CTF for manipulation,combine,transform and permutation some words or file text :p
A Deep Graph-based Toolbox for Fraud Detection
A library for making RepE control vectors
A modular, stack-agnostic toolkit of security review skills for AI coding agents to autonomously find, reproduce, and patch vulnerabilities.
AI gets the context. Not your private data. Local-first privacy proxy for browser chat, AI APIs, and coding agents.
Make your GenAI Apps Safe & Secure :rocket: Test & harden your system prompt
Raising the Cost of Malicious AI-Powered Image Editing
Benchmark of federated learning. Dedicated to the community. 🤗
Examples of techniques for training interpretable ML models, explaining ML models, and debugging ML models for accuracy, discrimination, and security.
Preimage attack against NeuralHash 💣
High-performance secrets scanner. CLI, Go library, Burp Suite extension, and Chrome extension. 487 detection rules with live credential validation.
Unified Multilingual Robustness Evaluation Toolkit for Natural Language Processing
Chrome extension which blocks requests to sites which have used legal threats to remove themselves from other blacklists.
Microsoft Security Copilot is a generative AI-powered security solution that helps increase the efficiency and capabilities of defenders to improve security outcomes at machine speed and scale, while remaining compliant to responsible AI principles
WinDBG Anti-RootKit Extension
GenAI compliance benchmark is a evaluation benchmarks for generative AI in regulated industries.
Find zero-days while you sleep. DeepZero is an automated vulnerability research framework that parses, decompiles, and analyzes thousands of Windows kernel drivers for exploitable IOCTLs natively using AI agents.
Real Intelligence Threat Analytics (RITA) is a framework for detecting command and control communication through network traffic analysis.
一款帮助云租户发现和测试云上风险、增强云上防护能力的综合性开源工具
Tools for understanding how transformer predictions are built layer-by-layer
A tool to transform Chromium browsers into a C2 Implant
Failure archive for ChatGPT and similar models
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
Code for IDS-ML: intrusion detection system development using machine learning algorithms (Decision tree, random forest, extra trees, XGBoost, stacking, k-means, Bayesian optimization..)
Parseltongue is a powerful prompt hacking tool/browser extension for real-time tokenization visualization and seamless text conversion, supporting formats like binary, base64, leetspeak, special characters, and multiple languages. Perfect for red teamers, developers, linguists, and latent space explorers.
Firefox and Chrome extensions to prevent Google from making links ugly.
VULNRΞPO - Free vulnerability report generator and repository, end-to-end encrypted! Templates of issues, CWE,CVE,MITRE ATT&CK,PCI DSS, import Nmap/Nessus/Burp/OpenVAS/Bugcrowd/Trivy, Jira export, TXT/JSON/MARKDOWN/HTML/DOCX, attachments, automatic changelog, stats, vulnerability management, bugbounty, local ai/llm, super fast pentest reporting!
Open-source adversary emulation for AI agents and MCP servers.
Casbin AI & MCP security gateway for HTTP, online demo: https://door.caswaf.com
Deliver safe & effective language models
A Privacy-Preserving Framework Based on TensorFlow
a browser extension to bring security and privacy to chrome, firefox, and opera
A Model for Natural Language Attack on Text Classification and Inference
🐘 🔒 KeePass-compatible browser extension for filling passwords.
Enjoy iCloud's Hide My Email service in your favourite browser
Official implementation of the paper "The Stable Signature Rooting Watermarks in Latent Diffusion Models"
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
PromptInject is a framework that assembles prompts in a modular fashion to provide a quantitative analysis of the robustness of LLMs to adversarial prompt attacks. 🏆 Best Paper Awards @ NeurIPS ML Safety Workshop 2022
Open-source runtime AI agent security tool - monitors and controls AI agents, catching malicious tool use, prompt injection, and policy drift in real time, before the agent acts.
Data-Driven Evaluation for LLM-Powered Applications
A robust, and flexible open source User & Entity Behavior Analytics (UEBA) framework used for Security Analytics. Developed with luv by Data Scientists & Security Analysts from the Cyber Security Industry. [BETA]
The official GitHub repository of the paper "Recent advances in large language model benchmarks against data contamination: From static to dynamic evaluation"
⚡ Vigil ⚡ Detect prompt injections, jailbreaks, and other potentially risky Large Language Model (LLM) inputs
Open-source AI penetration testing tool to find and fix your app’s vulnerabilities.
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
Source: GitHub API · Curated · Realtime