aifollow.news Search
Back Rohan Paul
Rohan Paul· @rohanpaul_ai · X· · Original publication time AI score54

John Hopkins and Carnegie Mellon University study: Small model learns to revise AI agent harness code from run resultsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI introduction

The editor reads an agent’s harness code and failure reports, then revises the code that controls what the agent sees and which tools it calls. Across 21 unseen reasoning task types, a 4B editor’s average edit score rose from 0.32 to 0.62, surpassing its 35B teacher; a separate QA test also showed transfer to other tasks.

Article · AI translation

约翰·霍普金斯大学和卡内基梅隆大学的一篇新论文表明,小模型可以根据运行结果学会改进智能体的运行框架代码,而且这种能力可以迁移到新任务。

一个经过训练、能根据失败报告重写智能体运行框架代码的小模型,可以帮助智能体适应新任务,所以不妨让它调优你的运行框架,而不是手动修改。

运行框架是决定模型看到什么以及调用哪些工具的代码。编辑模型读取运行框架和失败信息,然后编写代码改动,并根据新运行框架的得分获得奖励。

在21种未见过的推理任务类型上,一个4B编辑模型的平均编辑得分从0.32升至0.62,超过了它的35B教师模型。另一个在HotpotQA上训练的编辑模型,在另外2个问答基准测试中也能持续改进运行框架。

对每项任务运行多轮,因为反复编辑的效果优于单次修复。

来源:Rohan Paul · x.com

Research
Found an error? Send a correction