Bar charts and findings from LLM Agents Can Easily Tamper With Their Own Traces (arXiv:2609.30266v1), covering ten model–harness pairs.
0 comments
No comments yet.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 4 days ago
- Hacker News · 2 points · 4 days ago
- Hacker News · 1 points · about 13 hours ago
- Show HN: Groundtrack – Continual learning for coding agentsgroundtrack.devHacker News · 1 points · about 14 hours ago
- Hacker News · 45 points · 9 days ago
- Hacker News · 1 points · 10 days ago