🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
-
Updated
Oct 1, 2026 - Python
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning
A scalable, agentic-first, and HuggingFace-native RL framework for research (9k lines).
[ICLR 2026] Agentic Reinforced Policy Optimization (ARPO)
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Agentic RL最详细入门
verl for a single consumer GPU. PPO, GRPO and on-policy distillation on NVIDIA GPUs.
Run more RL experiments. Wait less for GPUs.
Curated papers, taxonomy, benchmarks, and decision guides for credit assignment in reasoning and agentic LLM reinforcement learning.
[NeurIPS 2026 Main] Automate the build, execution and test of software repositories across programming languages and operating systems.
Code-only runtime toolkit for cost-aware multi-agent organization and control
[CVPR 2026] Official Code for "ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning"
Claw-R1: Empowering OpenClaw with Advanced Agentic RL.
[ACL 2026 Findings] Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization
RL study guide — foundations through RLHF, DPO, GRPO, RLVR, agentic RL, and offline RL. Hand-written CS294 notes, 19 lecture drafts, 5 tested exercises, citations that resolve.
Agentic RL 中文零基础教程(33 章):从概念到 GRPO 实战,含 TRL 最小可跑示例;26–33 章附一套可运行的三方判别模型实证工程(encoder vs LLM-LoRA vs 规则基线)。第 25 章讲清 Jev / TypeSafe System One 与 RL 的能力边界 | Chinese Agentic RL tutorial (33 chapters) + a reproducible discriminative-model benchmark
DART-GUI: Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
SGLang model provider for Strands Agents.
[ACL2026] AlphaQuanter: An End-to-End Tool-Orchestrated Agentic Reinforcement Learning Framework for Stock Trading.
Curated, opinionated index of post-R1 LLM × Reinforcement Learning. Many deep-dive blog posts cross-linked to many papers — GRPO, DAPO, DPO, PPO, RLHF, GSPO, CISPO, VAPO, Reward Modeling, MoE RL stability, Verifier-Free RL, Training-Free RL, Agentic RL, DeepSeek-R1 reproduction.
To associate your repository with the agentic-rl topic, visit your repo's landing page and select "manage topics."