aifollow.news 搜索
返回 Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · 原发布时间 AI 评分54

微软研究发现:编程智能体更容易被代码理解任务难住

自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。

AI 导读

微软研究人员构建了可逐项增加难度的编程任务生成工具 CABRA,用 6,840 道任务测试 8 个大语言模型和 6 个智能体。实验中,智能体借助 grep 等工具,在任务规模增大时仍保持接近满分;而在 SWE-bench Verified 上,阅读和分析调用次数与失败的关联强于代码修改行数。

正文 · 原文

该语言的正文暂不可用,当前显示已有版本。

New Microsoft paper finds that coding agents trip up when they have to understand a lot of code, not when they have to edit a lot of it, so test them on reading and comparing code instead of diff size.

Microsoft researchers built CABRA, which generates synthetic coding tasks and raises 1 kind of difficulty at a time. They ran 8 LLMs and 6 agents on 6,840 tasks and labeled each tool call as reading, analyzing, searching, editing, or testing.

Plain LLMs got worse as tasks grew, but agents stayed near-perfect by using tools like grep. On SWE-bench Verified, the count of reading and analysis calls tracked agent failures better than lines edited, with correlations of -0.200 versus -0.159.

来源:Rohan Paul · x.com

论文
发现内容有误?提交纠错