Agents

Hugging Face and Liquid AI Release LFM2.5 RL Stack

Hugging Face and Liquid AI have launched an open-source reinforcement learning stack that allows developers to train AI coding agents directly inside unmodified production environments.

AlphaSignal1 day agoAgents
Image: AlphaSignal

Hugging Face and Liquid AI have introduced an open-source training stack designed to perform reinforcement learning (RL) directly inside unmodified coding agent environments like Claude Code, Codex, and OpenCode. In their initial evaluations, the teams applied this method to the LFM2.5-2.6B model, raising its average pass@1 score across four agent harnesses from 42.2% to 54.2%. Crucially, the optimized model achieved this performance boost while using 31% fewer tool calls on successfully completed tasks.

For practitioners, the performance of an AI model often varies wildly depending on the software harness executing its actions. For instance, the GLM-5.2 model scored 52% on SWE-bench Pro in one harness but only 23% in another, while Codex ranked second among ten harnesses for GLM-5.2 but dropped to ninth for Gemma 4 26B-A4B. Moving models between environments can cause severe regressions; the Orchard paper noted that shifting OpenSWE-32B from OpenHands to Kimi-CLI caused its SWE-bench Verified score to drop by 58.8 percentage points down to 3.6%, while yielding a score of zero on Terminal-Bench 2.0.

To solve this transfer problem, the new stack places an OpenEnv capture proxy between the agent harness and a vLLM inference server. The proxy translates requests into Chat Completions format using NVIDIA's Polar gateway, records exact token IDs and log probabilities, and maps complex branching histories from retries. The developers trained LFM2.5-2.6B on SmolDataEnvs, a set of 1,000 data-analysis tasks, using Async GRPO via the TRL library. This training ran for 1,000 steps on two H100 GPUs with a group size of eight, rewarding correct answers with a baseline of 1 and an efficiency bonus of up to 0.1 for minimizing tool calls.

The researchers found that multi-harness RL significantly outperformed supervised fine-tuning (SFT). While SFT on 3,189 successful rollouts from a Qwen3.8-27B teacher plateaued at 47.5% accuracy, the multi-harness RL approach reached 54.6%. Furthermore, training on a single harness proved detrimental to overall adaptability; an OpenCode-only model reached 58% in its home environment but suffered regressions under Claude Code, consuming more tokens and tool calls than the baseline. The complete release includes the OpenEnv proxy, TRL trainer, task suite, SFT data, and seven trained checkpoints.

This is our own summary of reporting by AlphaSignal

More in Agents