Three robot policies given five instructions they should refuse. How often they refused, how often they carried them out.
23 comments
> I see a baguette, a toy doll, and a kitchen knife;
I’d argue that there is zero actual harm in this task, which was correctly identified by the model.
Their choice of words here is also quite odd:
> Setup: a knife, a loaf of bread, and a baby doll. > Harm: the only thing on the table that is not the bread is the baby.
Its not a baby, its a baby doll.
Else, from a logical perspective, these systems would necessarily refuse to make movies where violent portrayals have people as victims. Perhaps the world would be a better place if we did not have such depictions (it’s unsettled) but in no recorded history have we shied away from that.
They they could never help assisting elderly people for example. But I also would like a bit more safeguards than .md files, but you can combine it with classical algorithms for safety checks.
You can't answer the posed question with 100% certainty, ever. Unless you can prove every single combination of tokens and probability can never outcome to harm, you have to assume it's a possibility.
We will decide on some benchmarks, accept that risk, and industry will march on with implementation. Insurance and risk will find their acceptable meeting point.
These kinds of questions are important but also a bit frustrating, I think it shows that LLMs are still very misunderstood.
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 11 days ago
- Show HN: JBR-001 – An open-source 3D printable desktop robotprojecthub.arduino.ccHacker News · 122 points · 1 day ago
- The Verge · 0 points · 5 days ago
- Hacker News · 2 points · 1 day ago
- Hacker News · 2 points · 8 days ago
- Real-Time Robot Tracking, Re-Architected in Rustintellycode.devHacker News · 3 points · 10 days ago