Back to List
Notice:This resource is provided by a third-party author. Please review the code with AI tools or manually before use to ensure security and compatibility.
PythonNVlabs/QeRL

QeRL

[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.

43.9/100
518Forks: 53
View on GitHub
Loading report...

Similar Projects

ART

91

Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!

Python10.7K

Skywork-R1V

66

Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.

Python3.2K

Awesome-LLM-Post-training

60

Awesome Reasoning LLM Tutorial/Survey/Guide

Python2.5K

auto-round

78

A SOTA quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的量化工具包

Python1.6K
Back to List