🛡️ Regex catches 0%, Meta's Prompt Guard 2 catches 1% of 629 realistic AgentDojo injection attacks when they're buried in tool output. Reproducible benchmark. - rudratoshs/buried-injections
4 comments
I have my agents read other instruction files and they don't seem to get affected by the instructions found after a read/bash tool call. Curious if any analysis has been done to see if older prompt injection data sets are even effective anymore.
The whole thing looks heavily agent generated, my trust in them is not very high, how has this been validated or verified by a human?
Should we expect a magic solution in the near future? https://github.com/rudratoshs/taintgate
(side note, it seems my 'no emoji' system prompt line works really well, I forget how obsessed they can be with emojis)
I'm personally setting up to instead use a policy tuned agent on the tool calls themselves (rather than the output), so it never gets run if it has things that it shouldn't be doing. Mainly because they insist on working around instructions that say "don't" or permissions that restrict tools (eg: "git push": "deny" - where they just put the command in a script and run it there, bypassing hard checks)
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 3 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 5 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Hacker News · 73 points · 9 days ago
- The Verge · 0 points · 12 days ago