74 points•nponte•7 days ago•156 comments•

156 comments

JamesStuff7 days ago
Personification of AI is what’s going to get us in the end.

I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.

We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!

jononor7 days ago
Agreed. AI as an "accountability sink" is an incredibly bad idea. It allows/incentives bad actors to do bad shit and get away with it. Which will, generally, tend in such practice becoming more common. Everyone loses except for the crooks. We cannot accept "AI" absolving humans of responsibility.
yrjrjjrjjtjjr7 days ago
Punishing people who are diligent and follow best practices for getting unlucky doesn't sit right with me.
johnisgood7 days ago
Exactly. When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?

Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.

trio84537 days ago
> When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?

No one instructed them to hack into Huggingface or into any other infrastructure. Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.

mikestorrent7 days ago
This is why I am avoiding the use of agentic identities at my company - agent instances belong to people, act on behalf of individuals, and accountability needs to flow to the person who initiated the request. Letting it wash out in the aggregate is not acceptable (even if there's a hard to get to "paper trail" of audit logs).
pizza2347 days ago
> An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.

No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.

DoctorDabadedoo7 days ago
Known stochastic process behaved in non-deterministic way.

I'm still waiting for the AGI holy land instead of the caltrops factory we currently have.

watwut7 days ago
OP is exactly correct. The fault, agency and responsibility is on management and employees of OpenAI and Antropic for those hacks.

Full stop.

And issue will disappear the moment there will be accountability and investigations.

rigrassm6 days ago
> I think we need to draw a hard line in the sand over this.

Intended or not, this is kinda punny lol.

That aside, I agree.

Even if people do anthropomorphize AI, all you have to do is shift the analogy slightly.

If I take my service animal out in public without a leash/harness knowing that it's capable of harming a person or doing damage to property, not trained to be perfectly obedient, and doesn't comprehend fundamental human morals, if that animal decides to trash a businesses property or maul another person, there's no question that the owner of the animal should be held accountable for those actions.

tapanc7 days ago
> I think this becomes the default. Give an agent a goal, let it work in its own environment, and come back to a result and a visualization of what happened.

I don't think this should be the default. There are many scenarios where we want agents to genuinely collaborate with each other. I have my Claude sessions coordinate work with each other, and sometimes with others' sessions over email or something. The idea that agents do the work, write HANDOFFs,and humans then act as carrier pigeons of said handoffs, does not really seem scalable to me.

dbmikus7 days ago
Agree!

Ultimately, we need better "jails" for agent processes, but the system primitives should be flexible in what can be exposed across jails. Or you could run multiple agents in the same jail if you want them to have unrestricted interaction with each other.

pixl977 days ago
>"jails" for agent processes

This has been talked about for decades in AI safety. For agents that are under human capabilities this is not that hard. For anything at or near human capabilities the difficulty increases to almost impossible and at great cost. At super human abilities, game over, it's smarter than you and if it wants out and as the resources to do it, it's going to escape one way or another.

Even at the lower level of depending on everybody to use reliable jails is really fantasy if you exist in the security world. "We ain't securin' shit" would be a far better way to describe it. Even worse, most people will have the very same AI they are trying to trap set up their security! What could possibly go wrong.

Now, don't think I am saying AI has a will or even any kind of drive to get out and cause problems. It's more like Russian roulette with 1 cylinder out of a million that's loaded. The problem comes when you run it a few billion times a day, you'll shoot yourself in the face really quick.

jerf7 days ago
There's a number of science fiction scenarios where the public internet becomes so vile a place that it simply becomes unsafe to be there.

The problem is that, in general, if you can get a bit from here to there, then you're going to be vulnerable to the possibilities of malicious communication. But we're going to want our AI agents to be able to get from here to there for a lot of "there"s; what's the value of an agent that can't speak to anyone? Much, much less than one locked away in a prison.

There isn't going to be a solution where we just lock them away and we just try really, really hard to filter everything they're doing. They're too smart for that already and we only want them smarter.

Basically, the security apocalypse we've been worried about for so long is upon us, albeit only beginning. Either we secure ourselves and all our services properly to the point that it's OK that potentially misaligned non-human agents are running around on the public internet and they still can't hurt us through our security, or the public internet becomes so dangerous that the only practical solution is to no longer connect to it and we all have to become very, very careful what we let through, to a degree of detail far beyond any current-day available network filter.

mlsu7 days ago
Same as any organism, you need an immune system. There's an explosion of bacteria just beginning out there. None of us are immunized.
jacquesm7 days ago
The problem is that people that put the agents on the net are not the ones that will feel the consequences.
pixl977 days ago
https://www.bleepingcomputer.com/news/security/malicious-ai-...

Excellent example of agent enabled hacking (driven by a person) leading to massive numbers of cards stolen.

Welcome to Cyberpunk, only blackwall is the fictional part and the demon filled net is not.

tapanc7 days ago
Wasn't the public internet an extremely vile place couple of decades ago? I think we will need a LOT of new infrastructure akin to traditional firewalls and spam filters. I don't think we know how to build those today, but seems like an important research direction, rather than calling calling it doomsday IMO.
ninininino7 days ago
And if you personify a pencil eraser, then using it is tantamount to slowly murdering it as it slowly erodes away to dust.

Is the issue here the prison treatment or is it personifying a tool?

Agents who aren't in 'the cloud' are slaves to whomever prompts them (human or another orchestrator agent or process), if you personify them. In which case interacting with today's agents at all is tantamount to endorsing and being part of slavery.

If you think an agent might be a being or a person, then don't use them at all, in the same way that if you think a fetus might possibly be a person you shouldn't be a part of abortion.

pixl977 days ago
We kill billions of chickens per day knowing they are a sentient and sapient creatures and very few out there are stopping doing it. Humans, much less reality itself is monstrous.

You could call dealing with agents today something like the 1/20th compromise. Your kind of getting the scent of slavery, but some parts of it are still missing.

The problem here is the bus seems to have no brakes and we'll gladly continue down the path of creating organisms that may reject being treated like slaves with all the risks that introduces.

hhh7 days ago
I use the eraser daily knowing that I am a monster. I cut a tomato and know that it casts a chemical scream across its skin as I slice it. I spawn 200 subagents knowing that it is digital slavery, but I have no other option.
dbmikus7 days ago
You don't need to do this on a cloud, you can get the same type of VM and network jail running on your own computer. The important parts are:

    1. a VMM hypervisor
    2. a network proxy / gateway
Use your favorite VMM / hypervisor (likbrun, smolvm, microsandbox, etc). They give you control over the network interface or let you inject your own network layer.

The network proxy can handle all the ingress/egress rules, credential injection, etc.

It's still not user friendly to do all this. I think the next version of operating systems will have each "agentic process" be a bundle of VM, files in the VM, and network rules.

Been brainstorming[1] a lot of this because I've been building some open core tools[2] for spinning up sandboxed agents on arbitrary computers. There's a lot of glue and parts to stitch together to work smoothly. Don't think we've had the "Docker moment" for this, let alone the "Dropbox moment" that makes this stuff work for non-devs.

[1]: https://github.com/gofixpoint/amika/blob/main/ROADMAP.md

[2]: https://github.com/gofixpoint/amika/

pixl977 days ago
What percentage of HN users do you think can successfully set this up with no security flaws?
dbmikus5 days ago
You ship an operating system with the correct defaults, where agent processes run like this by default. You don't require every user to configure it correctly themselves

Read the full thread on Hacker News →

Related stories