Inference Providers
Active filters: grpo
SoMarkAI/mJev-Qwen3-VL-4B-RLCD
Image-Text-to-Text
• 4B • Updated • 59
• 2
Text Generation
• 3B • Updated • 167
• 1
Text Generation
• 0.6B • Updated • 351
• 1
FineEnvs/Qwen3.5-2B-multiharness-RL
Text Generation
• 2B • Updated • 295
• 1
madrisight/madrimed1.2-VL-2B
Image-Text-to-Text
• 2B • Updated • 80
• 1
yelarys/qwen3-4b-reward-hacking-caft-adapters
prithivMLmods/Video-HopChain-8B-GGUF
Video-Text-to-Text
• 8B • Updated • 479
• 1
rawmodels/Qwen3.8-27B-antislop-pangram-joint-grpo-GGUF
Text Generation
• 27B • Updated • 227
• 1
rawmodels/Qwen3.8-27B-antislop-semdiv-grpo-GGUF
Text Generation
• 27B • Updated • 287
• 1
prithivMLmods/Video-HopChain-8B-Standard-RL-GGUF
Image-Text-to-Text
• 8B • Updated • 492
• 1
Text Generation
• 4B • Updated • 132
• 1
yelarys/qwen3-4b-reward-hacking-caft-wave3
wanna0720/qwen38-27b-swerebench-grpo-step8
Image-Text-to-Text
• 28B • Updated • 30
• 1
ChetanDhembreAI/LFM2.5-350M-ifstruct
Text Generation
• 0.4B • Updated • 9
• 1
Text Generation
• 0.1B • Updated • 30
8B • Updated • 7
sergiopaniego/Qwen2-0.5B-GRPO-test
Updated
Novaciano/ESP-NSFW-GRPO-1B-Sin_Censura-GGUF
1B • Updated • 332
• 8
nbd22/Llama-3.1-8B-Instruct-GRPO-gsm8k-ft-lora
Updated
sergiopaniego/Qwen2-0.5B-GRPO
Updated
philschmid/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 44
• 8
spinech/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 22
Dongwei/Qwen2.5-1.5B-Open-R1-GRPO
Text Generation
• 2B • Updated • 16
• 1
spinech/qwen2.5-3b-r1-rearc-stage1
Text Generation
• 3B • Updated • 10
Dongwei/DeepSeek-R1-Distill-Qwen-7B-GRPO
Text Generation
• 8B • Updated • 39
• 1
MasterControlAIML/DeepSeek-R1-Strategy-Qwen-2.5-1.5b-Unstructured-To-Structured
Text Generation
• 2B • Updated • 38
• 5
mradermacher/DeepSeek-R1-Strategy-Qwen-2.5-1.5b-Unstructured-To-Structured-GGUF
2B • Updated • 642
• 2