O
Online_RLHF
ZinYY/Online_RLHF
[NeurIPS 2025] A PyTorch implementation of the paper "Provably Efficient Online RLHF with One-Pass Reward Modeling". This repository provides a flexible and modular approach to Online Reinforcement Learning from Human Feedback (Online RLHF).
★94stars
Python
Updated: 2w ago
📋 Project at a Glance
Tap to expand
What's this?A open-source Security & Governance project, built with Python, focusing on large-language-model and llm
Who made it?Maintained by ZinYY team, 94⭐ on GitHub, #312 out of 413 in Security & Governance
Why does it exist?The ZinYY team recognized that existing large-language-model tools in Security & Governance were hard to use. Online_RLHF was designed to make llm more accessible.
What can it do?Key use cases: post-training, rlhf
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/ZinYY/Online_RLHF | 官网 https://arxiv.org/abs/2502.07193
🔗 github.com/ZinYY/Online_RLHF | 官网 https://arxiv.org/abs/2502.07193
Topics
large-language-modelllmpost-trainingrlhf