A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime…

2 points•sbulaev•6 days ago•0 comments•

0 comments

No comments yet.

Related stories