Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also…

3 points•potent_latent•8 days ago•0 comments•

0 comments

No comments yet.

Related stories