aifollow.news Search
Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score54

SkillGym trains AI agents on verified runs of human-written skillsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

The Shanghai AI Laboratory research turns human-written skill files into sandboxed tasks with code checkers, then trains a model on runs that pass. In Claude Code, Qwen3.5-35B-A3B’s Terminal-Bench 2.1 success rate rose from 39.33% to 58.43%. Without skill files, it scored 26.81% on SkillsBench, compared with 23.34% for the base model with skills.

Article · Original

The article text is unavailable in this language; an existing version is shown.

– https://t.co/dr1jJMbMEy

Title: "SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving"

ReplyRohan Paul@rohanpaul_ai
New Shanghai AI Laboratory paper finds that training on verified runs of human-written agent skills makes a model a better agent, even without the skill files. Prompt-time skills depend on retrieval and instruction following, and fine-tuning on verified skill runs reduces that dependence. Turning each skill file into a sandboxed task with a pass-or-fail checker produced training data that lifted Terminal-Bench 2.1 success by 19.10 points in Claude Code. Skill files usually sit in the prompt, so they only help if the agent finds and follows them. SkillGym turns each skill into a sandboxed task with a code checker, then trains on the runs that pass. In Claude Code, Qwen3.5-35B-A3B jumped from 39.33% to 58.43% on Terminal-Bench 2.1. With no skill files, it scored 26.81% on SkillsBench, beating the base model with skills at 23.34%. Loading the skills on top still helps, lifting it to 51.47%.
View replied-to post on X

来源:Rohan Paul · x.com

Research
Found an error? Send a correction