[ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)
PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL