aifollow.news Search
Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score54

Stanford paper distinguishes agent failures that call for harness changes from those that call for weight trainingMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

The study separates process failures, such as loops and exhausted step budgets, from content failures that produce poor plans. On a travel-planning benchmark, changes to the prompts, tools and checks around the model raised Qwen3.5-4B's score on held-out tasks from 0.16 to 0.30. Harness edits did not reduce the share of poor plans, while a LoRA adapter cut that share from 28% to 5% of Qwen3.5-9B's held-out runs.

Original post · Original

The article text is unavailable in this language; an existing version is shown.

@rohanpaul_ai

– https://t.co/GIQfgRAlFh

Title: "Harness Evolution Hits a Ceiling: When Weight Training Should Begin"

ReplyRohan Paul@rohanpaul_ai
New Stanford paper finds that harness changes fix agents that loop or stall, while agents that deliver bad plans need weight training instead. An agent can be improved by editing its harness, the prompts, tools, and checks around the model, or by fine-tuning its weights. They sorted failed runs into process failures, such as loops and used-up step budgets, and content failures, where a poor plan was delivered. On a travel-planning benchmark, an LLM-driven loop rewrote the harness, and its best runs then fine-tuned the model. Harness evolution lifted Qwen3.5-4B from 0.16 to 0.30 on held-out tasks, as plan delivery rose from 55% to 90%. Harness edits never shrank the share of poor plans, but a LoRA adapter cut them from 28% to 5% of Qwen3.5-9B's held-out runs. Before improving an agent, label why its runs fail, then fix process failures in the harness and content failures in the weights.
View this post on X ↗

来源:Rohan Paul · x.com

Research
Found an error? Send a correction