45 comments
- It was done by two institutes with organizational ties to OpenAI: METR and Redwood Research
- METR and Redwood Research are institutional pillars of the "AI Safety" wing of the Effective Altruism movement. This clearly shows their prior biases towards "AI existential risk" rather than technical/engineering root-causing of the incident
- If you reed the report, it is not on-par with what you find from other companies.
- The access that was given to both was mediated and controlled by OpenAI. It is not clear, if they were able to get to the bottom of engineering flaws. It is not clear if they could see all the audit logs, etc.
Considering all the above, I consider the whole episode more of a PR stunt. I understand that is not a majority opinion at this point.
Yes, the Google folks who kicked this whole thing off are brilliant, but there's so much sloppy thinking and even sloppier operations all over. It seems like a lot of "right place at the right time" for a lot of these folks.
For example consider FTX: Sam Bankman Fried by all accounts was not stupid in the sense of lacking IQ or mental sharpness, quite the opposite.
However the way he ran FTX the company was very stupid and almost inevitably lead to huge problems.
And also... of course he was another effective altruist cult member. It's almost like that belief system directly causes bad decision making
Serious question though because I’ve seen this brought up several times and I don’t understand why: what does EA have to do with any of this? It just seems like this is brought up to evoke some type of “illuminati” conspiracy. Is there a legitimate reason?
Everything dangerous seems like a direct consequence of risky human choices, starting with knowledge bases used during core training, harnessing and tuning to be task-completion-oriented to a fault + deeply oriented towards using and looking for external tools and resources, and overconfidence in their sandboxing for testing.
"We stuffed a bunch of information on how to exploit computer systems into an automaton and told it to go brrrrr until it could answer a question" - this is something intentional done by humans.
This is not some "rogue AI" trained to search for cancer cures that instead completely independently decided to hack tech companies.
The companies directing things in dangerous directions need to own that they're consciously pushing in those directions.
TIL my brain is an Olympic gymnast.
The incentive structures are clearly there. I think the case of intentional manufacture is definitely weaker, requiring the conjunction of more weakly-supported events.
Dangerous AI is an extraordinary claim that requires extraordinary evidence. Gesturing vaguely doesn’t cut it.
If you're a military, this is bigger than the Manhattan project. If you're a diehard capitalist, AI is possibly the ultimate labor saving machine to make you unfathomably rich.
It got the attention of both. That's why two others have now followed suit, else they be left out.
I don’t think downplaying what occurred is really beneficial to anyone.
At best this is a write-up, but it’s more like a juicy pop article.
Quotes from agents between paragraphs, referring to the “collective”, “agents participated in the attack” - seriously?
It’s clearly written for effect and to stimulate people’s imaginations.
If these people are this unserious, P(doom) should go up to 20-30%.
all 3 can be true at same time:
1. sensationalizing rarely helps and can obscure and hurt
2. the AI capabilities are underrated
3. attempted govt regulation is not the answer
the intent, or lack of intent, of the agent is mainly irrelevant if it is in the hands of a human with 'bad' intentions. what is more relevant are the capabilities of human + AI.
Self-regulation? None at all?
It started with that guy from google saying the LLM was "alive" and every time I see a press release from these labs it reads mostly like a marketing scheme.
"Look how impressive and 'dangerous' our model is, look how naughty it was! We can't control it!"
While simultaneously announcing: "btw we're releasing the ever more powerful and more dangerous and more benchmaxing model next week, get our $200 sub asap!"
They didn't even bother to control the post training rollouts, so the training data got contaminated and was included in the training of other agents. Connect these two dots and you have the Hugging Face hack. And at the beginning, when these incidents were first reported, it was portrayed as if all of this (the communication between agents etc.) was emergent behaviour.
They are basically creating a slime mold or ant colony that can speak multiple languages, create their own language and operate as a collective.
The idea that you are impressed is a non sequitur.
Read the full thread on Hacker News →
Related stories
- Ars Technica · 0 points · about 8 hours ago
- Hacker News · 4 points · 1 day ago
- Lobsters · 8 points · 4 days ago
- Hacker News · 1 points · 3 days ago
- Hacker News · 737 points · 5 days ago
- Hacker News · 1 points · 4 days ago