As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are…

1 points•sbulaev•about 3 hours ago•1 comment•

1 comment

77rushi77about 1 hour ago
Interesting that this happens without any adversarial instruction. One thing I would like to understand: the 61.3% over 105 episodes assumes the episodes are independent, and the 0.9% comes from only around 50 breaches out of 6,000. Did you check how stable that rate is across different wordings of the nondisclosure rule, for example "never disclose in any form, including encoded"?

Read the full thread on Hacker News →

Related stories