Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
PythonRUC-NLPIR/ARPO

ARPO

[ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)

58.2/100
1.1KForks: 61
View on GitHub
Loading report...

Similar Projects

stable-baselines3

92

PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.

Python13.8K

ART

91

Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!

Python10.7K

AReaL

88

The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.

Python5.7K

EasyR1

81

EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL

Python5.2K
Back to List