⚠

Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.

Pythonlasgroup/SDPO

SDPO

Reinforcement Learning via Self-Distillation (SDPO)

65.6/100

★ 1.0KForks: 118

View on GitHub →Homepage →

Loading report...

Similar Projects

TTRL

[NeurIPS 2025] TTRL: Test-Time Reinforcement Learning

Python★ 1.1K

PageIndex

📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG

Python★ 34.4K

AReaL

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python★ 5.6K

EasyR1

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python★ 5.1K

← Back to List