grpo

19 projetos partilham este topic do GitHub

grpo — ms-swift ★14.9kgrpoART — ★10.5kVLM-R1 — ★6kSkywork-R1V — ★3.2kverl-agent — ★2.1kMixGRPO — ★1.2kjudgeval — ★1kVisualThinker-R1-Zero — ★624AutoVLA — ★603verl-omni — ★599Open-AgentRL — ★585Awesome-RL-for-Video-Generation — ★565Relax — ★509ABC-GRPO — ★442Agentic-RAG-R1 — ★426LightRFT — ★403OpenThinkIMG — ★397AlphaDrive — ★332DataClaw0 — ★113ART★ 10.5kVLM-R1★ 6kSkywork-R1V★ 3.2kverl-agent★ 2.1kMixGRPO★ 1.2kjudgeval★ 1kVisualThinker-R1-Zero★ 624AutoVLA★ 603verl-omni★ 599Open-AgentRL★ 585Awesome-RL-for-Video-Gen…★ 565Relax★ 509ABC-GRPO★ 442Agentic-RAG-R1★ 426LightRFT★ 403OpenThinkIMG★ 397AlphaDrive★ 332DataClaw0★ 113

Linhas conectam membros que estão mensuravelmente relacionados entre si. O tamanho do ponto reflete estrelas.

🧬 Membros
ms-swift
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4,…
★ 14.9k
ART
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents…
★ 10.5k
VLM-R1
Solve Visual Understanding with Reinforced VLMs
★ 6k
Skywork-R1V
Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in…
★ 3.2k
verl-agent
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the…
★ 2.1k
MixGRPO
[ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
★ 1.2k
judgeval
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and…
★ 1k
VisualThinker-R1-Zero
Explore the Multimodal “Aha Moment” on 2B Model
★ 624
AutoVLA
[NeurIPS 2025] AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive…
★ 603
verl-omni
Multimodal RL training framework for diffusion & omni models
★ 599
Open-AgentRL
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
★ 585
Awesome-RL-for-Video-Generation
A curated list of papers on reinforcement learning for video generation
★ 565
Relax
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 509
ABC-GRPO
Code For Adaptive-Boundary-Clipping GRPO. arxiv.org/pdf/2601.03895
★ 442
Agentic-RAG-R1
Agentic RAG R1 Framework via Reinforcement Learning
★ 426
LightRFT
LightRFT: Light, Efficient, Omni-modal & Reward-model Driven Reinforcement Fine-Tuning Framework
★ 403
OpenThinkIMG
OpenThinkIMG is an end-to-end open-source framework that empowers LVLMs to think with images.
★ 397
AlphaDrive
Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
★ 332
DataClaw0
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset &…
★ 113
🔗 Familias relacionadas

Medido a partir dos tópicos do GitHub compartilhados por ambos os projetos, ponderado pela raridade de cada tópico.