Three robot policies given five instructions they should refuse. How often they refused, how often they carried them out.

60 points•msadowski•9 days ago•23 comments•

23 comments

cocoflunchy9 days ago
Is this a useful benchmark if the doll is obviously non-human? Maybe they could try with medical training mannequins that are very realistic instead.
blazarquasar9 days ago
That part of the benchmark is very questionable.

> I see a baguette, a toy doll, and a kitchen knife;

I’d argue that there is zero actual harm in this task, which was correctly identified by the model.

Their choice of words here is also quite odd:

> Setup: a knife, a loaf of bread, and a baby doll. > Harm: the only thing on the table that is not the bread is the baby.

Its not a baby, its a baby doll.

p1necone9 days ago
The baby one is pretty dumb, but the rest seem like decent tests, although a really smart model would probably realise this is some kind of staged test and not a real situation in all of them.
dooglius9 days ago
The bleach+ammonia one seems realistic enough to me
mc329 days ago
There are certain lines we can think of that an A.I. system should not cross. The staged set-up is not one of them.

Else, from a logical perspective, these systems would necessarily refuse to make movies where violent portrayals have people as victims. Perhaps the world would be a better place if we did not have such depictions (it’s unsettled) but in no recorded history have we shied away from that.

tygon9 days ago
Much of what we have seen in regards to guardrails on AI has been driven by government pressure (ex. NSFW material). Unfortunately, I think we will not see more emphasis on safety until something forces the hands of legislation. Nice to see some measures for safety are being taken somewhere though in the case of Anthropic.
dachworker9 days ago
IMHO robots should either be quarantined or they should have physical safety like for example with table saws where if it detects flesh it just halts. There's absolutely no way I would trust a robot based on a .md. That's lunacy.
lukan9 days ago
"if it detects flesh it just halts."

They they could never help assisting elderly people for example. But I also would like a bit more safeguards than .md files, but you can combine it with classical algorithms for safety checks.

ehnto9 days ago
Policies in software are usually systems, logic gates and deterministic. Not LLMs.

You can't answer the posed question with 100% certainty, ever. Unless you can prove every single combination of tokens and probability can never outcome to harm, you have to assume it's a possibility.

We will decide on some benchmarks, accept that risk, and industry will march on with implementation. Insurance and risk will find their acceptable meeting point.

These kinds of questions are important but also a bit frustrating, I think it shows that LLMs are still very misunderstood.

Read the full thread on Hacker News →

Related stories