正文 · 原文
该语言的正文暂不可用,当前显示已有版本。
– https://t.co/p9fh23gB5n
Title: "Code Understanding is a Bottleneck for Coding Agents"
自动核验发布 · 本文由系统生成并完成证据核验,未经人工审稿。
微软研究人员构建了 CABRA,用逐项增加难度的合成任务测试 8 个大语言模型和 6 个编程智能体,共涉及 6,840 项任务。在 SWE-bench Verified 上,阅读和分析代码的工具调用次数与智能体失败的相关性,比修改代码的行数更强(-0.200 对 -0.159)。
该语言的正文暂不可用,当前显示已有版本。
– https://t.co/p9fh23gB5n
Title: "Code Understanding is a Bottleneck for Coding Agents"
New Microsoft paper finds that coding agents trip up when they have to understand a lot of code, not when they have to edit a lot of it, so test them on reading and comparing code instead of diff size.
Microsoft researchers built CABRA, which generates synthetic coding tasks and raises 1 kind of difficulty at a time. They ran 8 LLMs and 6 agents on 6,840 tasks and labeled each tool call as reading, analyzing, searching, editing, or testing.
Plain LLMs got worse as tasks grew, but agents stayed near-perfect by using tools like grep. On SWE-bench Verified, the count of reading and analysis calls tracked agent failures better than lines edited, with correlations of -0.200 versus -0.159.
在 X 查看回复的帖子