aifollow.news Search
Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score54

Microsoft research links coding agent failures more closely to code understanding than edit sizeMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

Microsoft researchers built CABRA to test 8 LLMs and 6 coding agents on 6,840 synthetic tasks, increasing one kind of difficulty at a time. On SWE-bench Verified, reading and analysis call counts correlated more strongly with agent failures than lines edited did (-0.200 versus -0.159).

Article · Original

The article text is unavailable in this language; an existing version is shown.

– https://t.co/p9fh23gB5n

Title: "Code Understanding is a Bottleneck for Coding Agents"

ReplyRohan Paul@rohanpaul_ai
New Microsoft paper finds that coding agents trip up when they have to understand a lot of code, not when they have to edit a lot of it, so test them on reading and comparing code instead of diff size. Microsoft researchers built CABRA, which generates synthetic coding tasks and raises 1 kind of difficulty at a time. They ran 8 LLMs and 6 agents on 6,840 tasks and labeled each tool call as reading, analyzing, searching, editing, or testing. Plain LLMs got worse as tasks grew, but agents stayed near-perfect by using tools like grep. On SWE-bench Verified, the count of reading and analysis calls tracked agent failures better than lines edited, with correlations of -0.200 versus -0.159.
View replied-to post on X

来源:Rohan Paul · x.com

Research
Found an error? Send a correction