S
SDPO
lasgroup/SDPO
Reinforcement Learning via Self-Distillation (SDPO)
★1.0kstars
Python
Apache-2.0
Updated: 1d ago
📋 Project at a Glance
Tap to expand
What's this?Security & Governance project leveraging distillation and llm, built with Python, open-source
Who made it?Maintained by lasgroup team, 1K⭐ on GitHub, #115 out of 413 in Security & Governance
Why does it exist?The lasgroup team recognized that existing distillation tools in Security & Governance were hard to use. SDPO was designed to make llm more accessible.
What can it do?Key use cases: reasoning, rl
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/lasgroup/SDPO | 官网 https://self-distillation.github.io/SDPO
🔗 github.com/lasgroup/SDPO | 官网 https://self-distillation.github.io/SDPO
Topics
distillationllmreasoningrl