正文 · 原文
该语言的正文暂不可用,当前显示已有版本。
New Microsoft paper shows that agents can learn software skills from raw usage logs by keeping only skill edits that better predict users' next actions.
i.e. You do not need a live test environment to check whether a new agent skill helps, because next-action accuracy on old logs tracked live success.
Usage logs hold lots of know-how, but they record no goals, often mix several tasks, and cannot be replayed. Earlier methods, like Agent Workflow Memory, need goal-labeled examples or a live environment to test changes.
TeleTune guesses each session's goal and has the model predict every logged action using a text skill library. Wrong guesses suggest library edits, and an edit stays only if accuracy rises on held-out logs.
If your product records user activity, mine it for agent skills and judge each change by next-action accuracy on held-out logs.
– arxiv. org/abs/2610.05437
Title: "TeleTune: Evolving Agent Skills From Offline Telemetry"