Contribute to terrafying/ai-torture-chamber development by creating an account on GitHub.

1 points•rozumbrada•about 16 hours ago•1 comment•

1 comment

rozumbradaabout 16 hours ago
Steering language models into strong negative and positive valence states, and measuring what they say and what they're willing to do about it.

Read the full thread on Hacker News →

Related stories