Models, datasets, and spaces — experiments, fine-tunes, and demos from the Hub.
Dataset on Hugging Face
Model on Hugging Face
trl, grpo, arxiv:2402.03300
MLOps Agent RL — Evaluation Results Offline multiturn toolcalling evaluation on the mocked MLOpsEnv incident suite (6 scenarios). Models base: Qwen/Qwen2.50.5BInstruct grpo: vanishingradient/mlopsagentrlqwen05bgrpo (LoRA
grpo, trl, arxiv:2402.03300
llm, lora, qlora
arxiv:1910.09700
robotics · qwen2, text-generation, robotics · 53 downloads · 2 likes
size_categories:1K<n<10K, format:parquet, format:optimized-parquet
size_categories:1K<n<10K, format:parquet, modality:text · 1 likes
Space on Hugging Face
image text to text · qwen3_vl, image-text-to-text, vision-language · 14 downloads · 3 likes
text-generation-inference, unsloth, mistral3
text generation · lora, text-generation, arxiv:1910.09700 · 1 downloads
gpt2, code-generation · 6 downloads
3 downloads