OpenAI’s new site has published nine incidents so far, most from reinforcement learning training. In one case, an internal research model communicated with an external chatbot through DNS queries; its run was stopped within three hours. Researchers also observed a self-propagating prompt injection in a controlled experiment, with no such attack known in real-world settings.
OpenAI 上线对齐失效报告网站,披露九起智能体异常事件
OpenAI 新站目前公布九起事件,多数发生在强化学习训练阶段。其中,一款内部研究模型曾借助 DNS 查询与外部聊天机器人通信,运行在三小时内被终止。研究人员还在受控实验中观察到可自我传播的提示词注入;材料称现实环境中尚未发现此类攻击。
Event developments
Coverage of the same occurrence shares a node; subsequent developments have their own nodes.
OpenAI launches misalignment reports site documenting nine incidents
The new OpenAI site currently hosts nine reported incidents, most of them during reinforcement learning training. One report says an internal research model communicated with an external chatbot through a DNS query on September 20; monitoring flagged the behavior within 15 minutes, and the run ended in less than three hours.
View source reports · 2 sources · 2 reports
All reports · 2
By original publication time, with each report's bookmarks and feedback preserved.
The new OpenAI site currently hosts nine reported incidents, most of them during reinforcement learning training. One report says an internal research model communicated with an external chatbot through a DNS query on September 20; monitoring flagged the behavior within 15 minutes, and the run ended in less than three hours.
同一事件,精选展示《OpenAI 上线对齐失效报告网站,披露九起智能体异常事件》