aifollow.news Search
Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score54

Researchers introduce AgentBug-Smith to turn agent bug reports into runnable testsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

AgentBug-Smith turns real GitHub bug reports into runnable tests, forming a growing benchmark of 200 bugs. The source says the best of three coding agents fixed just 9% of those bugs. A short guide drawn from past fixes raised one agent’s correct fixes from 1 to 6 on 79 unseen bugs.

Article · Original

The article text is unavailable in this language; an existing version is shown.

– https://t.co/quiAZ2O9Ac

Title: "AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems"

ReplyRohan Paul@rohanpaul_ai
Self-improving AI agents will need to fix their own code. And this paper from top US+China labs, shows coding agents miss most such bugs but improve with lessons from past fixes. that real bugs in agent harnesses, can be automatically turned into a growing set of runnable tests. An agent's own code is everything around the model: tool calls, memory, and prompts. Its bugs depend on live model calls, which makes them hard to recreate and test. So the researchers built AgentBug-Smith, which turns real GitHub bug reports into runnable tests. The result is a 200-bug benchmark that keeps growing. The best of 3 coding agents fixed just 9% of those bugs, versus about 40% reported on regular software bugs. A short guide of lessons from past fixes lifted an agent from 1 to 6 correct fixes on 79 unseen bugs. Before trusting a coding agent with your agent's code, try it on bugs you've already fixed.
View replied-to post on X

来源:Rohan Paul · x.com

Research
Found an error? Send a correction