Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation…

1 points•sbulaev•2 days ago•0 comments•

0 comments

No comments yet.

Related stories