AI & ML interests

Internal testing artifact mangement for trl library

Recent Activity

sergiopaniegoΒ 
posted an update 9 days ago
view post
Post
639
Something I really like when I study a subject is understanding its history, how it reached the point where it is today

I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words

This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale

https://huggingface.co/blog/sergiopaniego/agentic-rl-2026
  • 1 reply
Β·
sergiopaniegoΒ 
posted an update 14 days ago
view post
Post
251
we just released a new blog "Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv"

you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced

and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine

the loop:
- OpenCode owns its tool loop inside an OpenEnv sandbox
- an in-sandbox proxy records the real token ids + logprobs, per turn
- a hidden-test verifier scores the result, and that is the reward
- TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL

blog + runnable example: https://huggingface.co/blog/sergiopaniego/trl-openenv-harness-training
  • 3 replies
Β·
sergiopaniegoΒ 
posted an update 15 days ago
view post
Post
2590
LFM2.5-2.6B just dropped!

and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.

basically, a full agent training pipeline but compressed into 2.6B

base model β†’ SFT β†’ specialized teachers per domain (SFT + RLVR) β†’ on-policy distillation back into one student β†’ agentic RL

the two most interesting stages

β†’ MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution

β†’ agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box

this makes a 2.6B that beats much larger models on instruction following and tool use

SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)

β†’ model: LiquidAI/LFM2.5-2.6B
β†’ blog: https://www.liquid.ai/blog/lfm2-5-2-6b
β†’ live series: https://www.youtube.com/playlist?list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5
  • 3 replies
Β·
sergiopaniegoΒ 
posted an update 20 days ago
view post
Post
2629
Simon Willison (@simonw ) has asked every new model to draw a pelican riding a bicycle for some time now

you look at the drawing and you know. but there is no number, so nothing can train against it, no?

I turned this idea into an rl env in OpenEnv. now, you can eval any model against it, and train against it with TRL

read the details!πŸ€“

https://huggingface.co/blog/sergiopaniego/pelican-env-openenv
  • 2 replies
Β·
sergiopaniegoΒ 
posted an update 21 days ago
sergiopaniegoΒ 
posted an update 23 days ago
view post
Post
2904
quick reminder! 🚨

tomorrow (Tuesday, July 28), we're back with Class 3 of the Training Agents live series

🧠 what: reinforcement learning for training agents (GRPO): how it works, how to implement it in TRL, and end-to-end examples
πŸ—“οΈ when: Tuesday, July 28 - πŸ•” 5:00 PM CEST / 8:30 PM IST
πŸ“ where: Live on @huggingface 's X, YouTube, and LinkedIn

live: https://www.youtube.com/watch?v=ztdTed5egrM

class 1: https://x.com/SergioPaniego/status/2069382207618379813
class 2: https://x.com/SergioPaniego/status/2075180665184686187
  • 1 reply
Β·
sergiopaniegoΒ 
posted an update 26 days ago
view post
Post
227
you can now train your own coding agents with trl + openenv, starting with opencode

we just added end-to-end support for training agent harnesses:

> TRL: a loop-owning training path (AsyncGRPOTrainer + HarnessRolloutWorker) that launches the agent in an OpenEnv session, reads back its trace, reconstructs the training samples, and trains with AsyncGRPO
> OpenEnv: the OpenCode harness environment plus a transparent proxy that forwards the agent's model calls and records each turn's token ids and logprobs

you train the actual opencode agent as is, it runs its own loop and tools and the policy learns from the exact tokens it produced

we're shipping a self-contained example: local subprocess sandbox, DeepCoder problems, validated on Qwen3-8B.

> example: https://github.com/huggingface/trl/blob/main/examples/scripts/openenv/opencode.py
> docs: https://huggingface.co/docs/trl/main/openenv

and we're working actively on both sides so expect more πŸ€“
  • 1 reply
Β·
sergiopaniegoΒ 
posted an update 27 days ago