aifollow.news 搜索
返回 Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · 原发布时间 AI 评分54

斯坦福论文区分 AI 智能体两类失败:卡住可改运行框架,计划差需训练权重

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 导读

这项研究将智能体失败分为循环、耗尽步骤等流程问题,以及交出糟糕计划的内容问题。在旅行规划测试中,调整模型周围的提示、工具和检查后,Qwen3.5-4B 在留出任务上的得分从 0.16 升至 0.30;框架调整未减少糟糕计划的比例,而 LoRA 微调将 Qwen3.5-9B 留出运行中的该比例从 28% 降至 5%。

原帖引用 · 原文

该语言的正文暂不可用,当前显示已有版本。

@rohanpaul_ai

– https://t.co/GIQfgRAlFh

Title: "Harness Evolution Hits a Ceiling: When Weight Training Should Begin"

回复Rohan Paul@rohanpaul_ai
New Stanford paper finds that harness changes fix agents that loop or stall, while agents that deliver bad plans need weight training instead. An agent can be improved by editing its harness, the prompts, tools, and checks around the model, or by fine-tuning its weights. They sorted failed runs into process failures, such as loops and used-up step budgets, and content failures, where a poor plan was delivered. On a travel-planning benchmark, an LLM-driven loop rewrote the harness, and its best runs then fine-tuned the model. Harness evolution lifted Qwen3.5-4B from 0.16 to 0.30 on held-out tasks, as plan delivery rose from 55% to 90%. Harness edits never shrank the share of poor plans, but a LoRA adapter cut them from 28% to 5% of Qwen3.5-9B's held-out runs. Before improving an agent, label why its runs fail, then fix process failures in the harness and content failures in the weights.
在 X 查看这条帖子 ↗

来源:Rohan Paul · x.com

论文
发现内容有误?提交纠错