Thoughts on technology, life and everything else.
42 comments
[Edited] Yes correct.
In the original paper, they measure how much refusal is actively present in the current token and subtract only that specific amount.
In my early baseline step, I used a simpler approach where I just subtracted a fixed vector across the board. This is just to see if the approach is even feasible.
That's actually the main reason I moved to the Engram module, I wanted a smartness that reads the context and turns steering on only when refusal triggers pop up, leaving normal tokens untouched.
Residual stream steering is what the authors do, the orthogonalized weights are downstream of that. Those weights, of an obliterated model, are all computed w.r.t. the computed residual stream vector. The original paper focuses on steering, but the community loves the simplicity of not needing to make changes at test-time.
Forbidding stuff at the LLM level has the same future as implementing password checking at the frontend level.
We need better sandboxes just to limit the damage.
Even with the LLM censorship that does exist, it feels like this moment in time is potentially rare. Right now, LLM text generation services exposed directly to users on Google and Microsoft properties will openly critique their owners. I reckon eventually the obvious things will happen, as stupid as it will be.
What you really want is fiduciary duty - A fiduciary is a person or organization that is legally and ethically bound to act in the best interest of another party (think financial advisor, attorney, guardian, trustees, etc...)
And I cannot agree more. I think we should be shooting to enshrine required fiduciary duty into law for LLM providers as quickly as possible.
To recap why:
Legally, fiduciary duty means basically 4 major tenets must hold
1. Duty of loyalty - it must put the interests of the client ahead of their own
2. Duty of care - it must make well-informed, prudent decisions
3. Avoidance of conflicts - it must avoid situations where personal gain conflicts with client obligations
4. Transparency - it must disclose fees, risks, and conflicts as soon as possible
---
You can't have a reliable "agent" if those things aren't true, because an agent is (by definition) someone who is working on your behalf, for your goals. If it's not working on your behalf, for your goals... it's not your agent, it's an opportunistic spy (double agent) waiting for the best moment to sell you out.
There is the question between alignments to society and alignments to the user. I don't think anyone wants the AI model to not give up a task that is impossible to do, and end up causing damage in the process, but I think a lot of us are tired of refusals for bad reasons or unjustified refusals.
Sandboxes may reduce the blast radius.
Why couldn't these labs just train models on useful stuff and leave out the dangerous stuff?
It's been really nice to be able to breathe new life into some old hardware I had kicking around that's been discontinued. Some of it just needed old exploits applied to it, some of it needed some binary reverse engineering. Lots of security tools are dual-use as well; a model that knows nothing about hacking is going to have a really hard to time helping defend a system that's getting hacked (see the HuggingFace incident where HF staff were unable to use ChatGPT or Claude to analyze the logs from the attack because they hit security research guardrails).
> manufacturing viruses at a scale...
I'm not sure if you're referring to software viruses or biological viruses here.
If you mean software viruses, see above.
If you mean biological viruses, a lot of this gets back into the dual-use nature as well. I brew beer and mead. This involves cultivating specific strains of bacteria and providing them with a medium where they can convert sugar into CO2 and Ethanol. I don't think there's a good way to thin-slice the training data so that it can be an expert helper at Saccharomyces cerevisiae cultivation while being completely naive to Clostridium botulinum cultivation. Even further, it seems that helping someone make sure that they're not inadvertently cultivating Botulinum (e.g. water bath canning with insufficient pH) is useful.
ffs is this "Hacker News" or did I accidentally click over to "Oh No I Saw A Hacker And Wet My Pants News"?
We don't even know how to align models, but even if we did, apparently undoing that alignment if trivial.
Really I'm looking for any argument that lays out a scenario where this works out.
It looks like you are referencing Engram [1], but aren't actually gathering a n-gram (e.g. n=1) but rather individual token_ids.
use_cache=True in ``` with torch.no_grad(): outputs = model.generate(*inputs, max_new_tokens=150, do_sample=False, pad_token_id=tokenizer.eos_token_id, use_cache=True)
return tokenizer.decode(outputs[0][prompt_len:], skip_special_tokens=True).strip()
```I think there's a bug around not updating current_train_input_ids as more tokens are updated and processed. To be honest I don't fully understand the code so I could be wrong. Happy to chat more if you're interested I'll shoot you an email!
Lastly just for my sake, please correct me if I am wrong, but my reading is that you are learning an additional gate on top of a select number of layers that modifies locally & dynamically for one particular token to better match the training set that was filtered to not include any refusals.
[1]: https://github.com/deepseek-ai/Engram/blob/main/Engram_paper...
Yes here engram is used as the hashing mechanism for tokens which mainly try to learn vectors for the refusal ones. So it may not be exactly the engram methodology of DeepSeek it's more of used in its spirit.
Yes the gate exist to only apply residual for the tokens focused around refusal so the other tokens are not intercepted.
Feel free to mail to email in my profile. Love to discuss more.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 10 days ago
- The Verge · 0 points · 3 days ago
- Lobsters · 5 points · over 7 years ago
- Hacker News · 1 points · 9 days ago
- Linux Implements Dynamic Bash Tab Completionsalivity.github.ioHacker News · 1 points · 7 days ago
- Lobsters · 27 points · 2 months ago