Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also…
0 comments
No comments yet.
Related stories
- Hacker News · 3 points · 4 days ago
- Ars Technica · 0 points · 13 days ago
- Observations on LLMs for Privacy Workdesfontain.esHacker News · 3 points · 11 days ago
- Hacker News · 9 points · 11 days ago
- Hacker News · 1 points · 11 days ago
- Hacker News · 16 points · 12 days ago