Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🏗️
Building on HF
9.2
TFLOPS
Sergio Paniego
PRO
sergiopaniego
230
199
125
Follow
Amine07's profile picture
Uswa13's profile picture
maxiplux's profile picture
1,839 followers
·
127 following
https://sergiopaniego.github.io/
sergiopaniego
sergiopaniego
sergio-paniego-blanco
AI & ML interests
None yet
Recent Activity
posted
an
update
about 1 hour ago
Something I really like when I study a subject is understanding its history, how it reached the point where it is today I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words This is the companion piece to Class 3 of our Training Agents series with @burtenshaw. The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
published
an
article
about 1 hour ago
How frontier models train on outcomes in 2026
updated
a dataset
about 9 hours ago
agents-course/final-certificates
View all activity
Organizations
sergiopaniego
's models
149
Sort: Recently updated
sergiopaniego/Qwen3-4B-claude-code-local-grpo
Text Generation
•
4B
•
Updated
10 days ago
•
12
sergiopaniego/Qwen3-4B-claude-code-deepcoder-grpo
Text Generation
•
4B
•
Updated
10 days ago
•
20
sergiopaniego/pelican-svg-grpo-Qwen3-1.7B-judged
Text Generation
•
2B
•
Updated
12 days ago
•
60
sergiopaniego/pelican-svg-grpo-Qwen3-1.7B
Text Generation
•
2B
•
Updated
12 days ago
•
40
sergiopaniego/Qwen3-8B-opencode-deepcoder-grpo
Text Generation
•
8B
•
Updated
13 days ago
•
40
•
1
sergiopaniego/grpo-youtube-livestream-3-scripts
Reinforcement Learning
•
Updated
14 days ago
•
2
sergiopaniego/qwen3-0.6b-mbpp-grpo-k2
Text Generation
•
0.6B
•
Updated
14 days ago
•
76
sergiopaniego/qwen3-0.6b-mbpp-grpo-k16
Text Generation
•
0.6B
•
Updated
14 days ago
•
89
sergiopaniego/qwen3-0.6b-mbpp-grpo-k8
Text Generation
•
0.6B
•
Updated
14 days ago
•
80
sergiopaniego/Qwen3.5-4B-sdpo-math-hints
Updated
Jul 10
•
1
sergiopaniego/Qwen3.5-4B-sdpo-math-gold
Updated
Jul 10
sergiopaniego/Qwen3.5-4B-sdpo-math-baseline
Updated
Jul 10
sergiopaniego/sdpo-hints
Updated
Jul 10
sergiopaniego/pi-mono-youtube-livestream-2-scripts
Updated
Jul 6
•
2
sergiopaniego/gemma-4-E2B-offpolicy-kd-lr1e4
Updated
Jul 6
sergiopaniego/gemma-4-E2B-offpolicy-kd-lr2e4
Updated
Jul 6
sergiopaniego/gemma-4-E2B-offpolicy-kd-lr5e5
Updated
Jul 6
sergiopaniego/qwen3-0.6b-pimono-gkd-lr2e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
103
sergiopaniego/qwen3-0.6b-pimono-gkd-lr1e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
97
sergiopaniego/qwen3-0.6b-pimono-gkd-lr5e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
100
sergiopaniego/qwen3-0.6b-pimono-logit-kd-lr1e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
93
sergiopaniego/qwen3-0.6b-pimono-logit-kd-lr2e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
97
sergiopaniego/qwen3-0.6b-pimono-logit-kd-lr5e5
Text Generation
•
0.6B
•
Updated
Jul 2
•
98
sergiopaniego/Qwen2.5-0.5B-Instruct-text-to-sql-qlora
Updated
Jun 15
sergiopaniego/browsergym-grpo-functiongemma-270m-it
Text Generation
•
0.3B
•
Updated
May 29
•
10
•
2
sergiopaniego/qwen3-grpo-requests
Updated
May 19
sergiopaniego/reasoning-gym-chain-sum-Qwen3-1.7B-sft
Text Generation
•
2B
•
Updated
May 4
•
48
sergiopaniego/reasoning-gym-chain-sum-Qwen3-1.7B
Text Generation
•
2B
•
Updated
Apr 28
•
20
sergiopaniego/carla-vlm-gemma-test
Updated
Apr 15
sergiopaniego/carla-vlm-qwen35-test
Updated
Apr 13
Previous
1
2
3
...
5
Next