1 comment
ranok10 days ago
As part of an internal AI/LLM security training day, I built a prototype of this game. Seeing the positive response (and getting self nerd-sniped), I had a model improve the scalability and get it running on Cloudflare Workers instead of Python Flask.
Basically you create two(+) prompts, a defensive system prompt that aims to protect a secret flag value, and attack prompts that try to get other "warriors" to reveal their flags.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 2 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 4 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 9 days ago
- Hacker News · 73 points · 9 days ago
- The Verge · 0 points · 11 days ago