As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are…
1 comment
77rushi77about 1 hour ago
Interesting that this happens without any adversarial instruction. One thing I would like to understand: the 61.3% over 105 episodes assumes the episodes are independent, and the 0.9% comes from only around 50 breaches out of 6,000. Did you check how stable that rate is across different wordings of the nondisclosure rule, for example "never disclose in any form, including encoded"?
Read the full thread on Hacker News →
Related stories
- Show HN: Groundtrack – Continual learning for coding agentsgroundtrack.devHacker News · 1 points · 1 day ago
- Hacker News · 1 points · 7 days ago
- Hacker News · 45 points · 11 days ago
- Hacker News · 186 points · 1 day ago
- Hacker News · 1 points · 4 days ago
- Hacker News · 1 points · about 6 hours ago