A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime…
0 comments
No comments yet.
Related stories
- Hacker News · 55 points · 4 days ago
- Hacker News · 1 points · 8 days ago
- Hacker News · 1 points · 1 day ago
- Hacker News · 1 points · 7 days ago
- Hacker News · 1 points · 8 days ago
- Ars Technica · 0 points · 13 days ago