Evaluation & Alignment25 / 77 projects
Model alignment leaderboard featuring the hottest AI evaluation and alignment projects on GitHub for safe AI development.
OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.
Robust recipes to align language models with human and AI preferences
SWE-bench: Can Language Models Resolve Real-world Github Issues?
Align Anything: Training All-modality Model with Feedback
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard
A unified evaluation framework for large language models
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
Secrets of RLHF in Large Language Models Part I: PPO
OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI
The official GitHub repository of the paper "Recent advances in large language model benchmarks against data contamination: From static to dynamic evaluation"
The official evaluation suite and dynamic data release for MixEval.
BeaverTails is a collection of datasets designed to facilitate research on safety alignment in large language models (LLMs).
A benchmark suite for evaluating LLM-based interactive scientific reasoning.
Code Repository for: AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
LLM alignment jailbreak; a set of instructions for auditing their internal reasoning and uncovering biases
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost
A benchmark for evaluating LLM × harness performance.
(EMNLP 2025 Findings) Source Evaluation scripts for Humanity's Last Code Exam
Code repo for "LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners"
[NeurIPS 2025] A PyTorch implementation of the paper "Provably Efficient Online RLHF with One-Pass Reward Modeling". This repository provides a flexible and modular approach to Online Reinforcement Learning from Human Feedback (Online RLHF).
Multi-modal AI-generated content detection: image, video, and audio. Benchmarks, training code (DINOv2, DINOv3, ReStraV, BreathNet), and evaluation pipeline for real vs. synthetic classification with calibration-aware metrics.
This repo contains the code for "MEGA-Bench Scaling Multimodal Evaluation to over 500 Real-World Tasks" [ICLR 2025]
An benchmark for evaluating the capabilities of large vision-language models (LVLMs)
Source: GitHub API · Curated · Realtime