S
safe-rlhf
PKU-Alignment/safe-rlhf
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
★1.6kstars
Python
Apache-2.0
Updated: 1w ago
📋 Project at a Glance
Tap to expand
What's this?A , built with Python open-source project in the Security & Governance category, core strengths: ai-safety/alpaca
Who made it?Maintained by PKU-Alignment team, 1.6K⭐ on GitHub, #95 out of 413 in Security & Governance
Why does it exist?As the Security & Governance landscape evolved, the PKU-Alignment team identified the need for better ai-safety solutions. safe-rlhf was created to simplify alpaca workflows.
What can it do?Key use cases: beaver, datasets, deepspeed
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/PKU-Alignment/safe-rlhf | 官网 https://pku-beaver.github.io
🔗 github.com/PKU-Alignment/safe-rlhf | 官网 https://pku-beaver.github.io
Topics
ai-safetyalpacabeaverdatasetsdeepspeedgptlarge-language-modelsllamallmllmsreinforcement-learningreinforcement-learning-from-human-feedbackrlhfsafe-reinforcement-learningsafe-reinforcement-learning-from-human-feedbacksafe-rlhfsafetytransformertransformersvicuna