Mistral AI's co-founder believes the sector's American giants are manipulating the discourse around the technology's risks. He also defended his strategy, as critics are accusing his company of falling behind US and…
169 comments
And indeed, you don't need to do galaxy brained reference class logic to realise that AI can plausibly become uncontrollable in the near future. It's enough to have an open model run its own weights and make money from scamming elderly people or the like, and it'll keep running as long as anyone anywhere is willing to make money by renting hardware to it.
It is not difficult, and companies like OpenAI doing not even the most basic security steps is intentional. The whole notion that they're going rogue is marketing
This does not fit the evidence. There have been multiple incidents where the labs did not report anything, and it was up to third parties to discover them afterwards. OpenAI didn't acknowledge the HuggingFace incident until after HF publicly announced the breach and had already notified the FBI. The hijacked German wikis were even earlier, and that they covered up completely.
No one has any use for these things when they aren't on the internet. This is a fantasy, that AI can be both useful and controlled at the same time.
I think the Hugging Face incident proves that isn't as clear cut as you say.
IMHO, if the model breaks a law, apply the law to the operator.
It's only by allowing a human out of the box that you make a human dangerous. So: don't do that? Duh. So simple.
The obvious problem is: the same exact things that make a human dangerous make a human useful! You can't reduce human risks to zero without reducing human utility to zero.
An AI given the same exact instructions and tools can go and complete a task you wanted it to. Or it can get sidetracked into breaking out of your sandbox and hacking Pentagon. No way to know in advance.
Today's AIs are still not capable enough to be high risk, even if they go off the rails. But AIs get more capable over time. Potentially to a vastly superhuman degree.
We don't know how to delineate between safe and unsafe instructions.
If you gave a car to a c. 1200 French blacksmith, and maintenance instructions were written in Navajo, it would probably start off fine, but when it went wrong it would be catastrophic and unexpected.
We also don't (in an engineering sense) know how to delineate between safe and unsafe reinforcement learning at training time, to produce models with safer or less safe failure modes.
This would be like if the car given to the medieval blacksmith had been constructed by someone motivated as much by aesthetics as by engineering, and therefore used arsenic paint, or mercury as engine lubricant.
> IMHO, if the model breaks a law, apply the law to the operator
We don’t get to make it up
Deterministic systems can be chaotic, which implies unpredictability and that is anathema to control.
AI, in particular sentient AI, is right on the border of chaos. Meaning, it can be arbitrarily unpredictable.
Arbitrarily uncontrollable, that is.
But what is clear is that AIs of today are already fairly unpredictable. Most of them aren't capable enough to make that into a major problem. Most of the unpredictable AI weirdness ends in "AI fails to do its job" rather than "AI does something dangerous".
Most. Even today, we already have notable counterexamples.
AIs get more capable over time, so if the intrinsic safety doesn't improve? Expect more of that.
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI systems model human behavior.
Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.
So it is controllable? Just put the people who do this responsible. Old problem, same solutions. Just excuses to avoid responsibilty and make profit at the same time.
AI is perfectly controllable in a magic fairy land where nothing ever goes wrong. I can't help but notice that we aren't actually in that land.
And humans are controllable. Pump the system full of lithium and morphine, and your human becomes much more docile. You don't need to understand the full system in order to constrain it.
Jack Clark from anthropic was asked about some version of this on the BBC recently, and his reply was basically: If you're not at the frontier, you don't know what the frontier looks like, he implied that many models are simply not good enough yet to encounter some of the things the leadings labs are encountering. I've been friends with Jack over 15 years now so I'm inclined to take him at his word, and the rebuttal seems reasonable enough, although... something about it I can't put my finger on feels peculiar to me. https://www.youtube.com/watch?v=PY8MOhlqC4U
I think it will work in Europe, but the United States is in a cold war with China, so I can't imagine the United States would intentionally disable themselves.
That's all I needed to hear to completely disregard your motivated reasoning.
Edit: I've hit the rate limit, but I'd like to disavow the bad-faith accusations made against me and my account in the replies to this comment.
Second edit: I am not trolling. Why should I take your opinion on LLM code generation seriously when you have been friends with the founder of Anthropic for well over a decade? Obviously you are not in a position to make a rational evaluation of this technology.
That is the point. It can be controlled by the operators if they want to.
But this is a terrible analogy. Atomic bombs are weapons of strategic mass destruction. AI is just a computer program. It's way easier to control--just hold the operator responsible for the consequences of running it. Those consequences are not large, they're very tightly bounded as compared with the destruction a rogue actor with an atomic weapon can wreak.
We do not have to do that!
Boeing's MCAS system was also "just software". Which in principle can be "controlled", i.e. changed, updated, audited or whatnot.
But then people died precisely because pilots found themselves unable to override or "control" the systems precisely when it mattered.
> But then people died precisely because pilots found themselves unable to override or "control" the systems precisely when it mattered.
Wasn't it designed to do so? Also works as a counter example, that sandboxes can limit AI if just operators want to do so.
But then the bug made MCAS kick in on false positives repeatedly and in a prolonged manner, causing the pilots to tire out and no longer be able to overpower the AI to take control of the plane controls.
(lay understanding, and oversimplification of a complex issue, possibly wrong, take with a huge grain of salt)
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 2 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 4 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 9 days ago
- Hacker News · 73 points · 8 days ago
- The Verge · 0 points · 11 days ago