Qwen3-8B teacher-regularized RL method collection for the ToolUse dataset. Includes GRPO-TR, RLSD-TR, SDPO-TR, and SRPO-TR models.
-
SeongryongJung/Qwen3-8B-Tooluse-GRPO-TR
Text Generation • 8B • Updated • 8 -
SeongryongJung/Qwen3-8B-Tooluse-RLSD-TR
Text Generation • 8B • Updated • 34 -
SeongryongJung/Qwen3-8B-ToolUse-SDPO-TR
Reinforcement Learning • 8B • Updated • 5 -
SeongryongJung/Qwen3-8B-ToolUse-SRPO-TR
Reinforcement Learning • 8B • Updated • 5 • 1