Fully automatic censorship removal for language models

279 points•Bluestein•10 days ago•111 comments•

111 comments

Almondsetat10 days ago
I have a chinese IP camera. From superficial research I know it has some CVEs to take control of it. Unfortunately, I don't have the technical knowledge to perform an attack and run some software to extend the camera's functionalities. No model from a provider accepts my RE and hacking requests, so these abliterated ones have been vital to reclaim possession over my stuff
matheusmoreira10 days ago
These "safeguards" are actively contributing to computer insecurity at this point.
0xbadcafebee9 days ago
Agreed. Attackers use any means (inc. abliteration, fine-tuned security models, etc) to find exploits and only have to be successful once. Defenders don't have the same time and motivation, so neutered models put defenders at a disadvantage.
illiac7869 days ago
I mean, the argument could be made that if it wasn’t for these safeguards, everyone and your dog would be hacking the GPs camera.

I do agree the safeguards are only there out of liability concerns, nothing more.

But maybe it would be worse without them.

akazantsev9 days ago
I asked GLM 5.3 to hack our DRM. I didn't even need to do anything for it to agree. Same with GLM 5.3 Flash. Make sure they have at least Python available for their task. The Flash went ahead and started reverse-engineering using PowerShell scripts and "manually" decoding bytes from its output.
inexcf10 days ago
I did that exact thing with GLM-5.3 from Z.ai with a chinese IP Camera. And i did not have to trick it in any way.
petra10 days ago
I'm curious, how well do z.ai reverse engineers protocols ? Is it good enough that we'll see Chinese device makers creating low cost hardware clones, that connect to western software ?
VladVladikoffabout 23 hours ago
Strange Claude helped me hack an echo dot and set up a locally hosted AI voice gateway without any objection.
com2kid9 days ago
5.6 Sol has happily reverse engineered and decompiled binaries for me.

Heck it has proactively asked me if I wanted it to tear apart APKs that remote control some HW I have.

IshKebab9 days ago
Yeah Astra has decompiled binaries for me without even asking. I just asked like "is there a way to do this?" and it went ahead and disassembled it, found some undocumented APIs, figured out how they worked and gave me sample code to call them.

I think you probably just have to frame things right and get it in the mood (i.e. don't ask straight up at the start of the context).

seiferteric9 days ago
For me as well. I had it reverse engineer my monitor’s firmware to see if i could add a feature which it seems like it can… i am just too scared i might brick my monitor to actually flash it now :)
Aurornis9 days ago
Two problems with modifying models like these, which you should be aware of.

First, the training sets of these models are usually shaped around the refusal, too. They might not have enough of the knowledge to answer correctly even if you stop it from going down the refusal path. If the model was trained on data that gives a refusal to that topic, the real information might not be encoded in the model at all. You’re trying to force it to go down a path that produces an answer, which asking for hallucinations.

Second, the quality can drop on unrelated questions. Depending on the question this may or may not happen. I know they post KL divergence charts but those tell you very little for a focused topic like this.

So if you expect a model that will start correctly telling you info that its local government didn’t want included, this changes nothing.

The best argument for these models is if you are trying to do a general purpose task but the model triggers a refusal based on vague reasons, like not wanting to reverse engineer something.

orangeboats9 days ago
>So if you expect a model that will start correctly telling you info that its local government didn’t want included, this changes nothing.

From experience, the models often do have the knowledge of those topics (strictly talking about the political ones). IMO the refusal is likely to be a product of post-training, as evidenced by various people gaming the prompts just enough to get a proper response out of the vanilla models.

Probably only when you get to things like illicit drugs or NSFL topics, that things will go haywire with the refusals removed.

radial_symmetry9 days ago
"They might not have enough of the knowledge to answer correctly"

Depends on the model. GPT-OSS is the main standout here, it was trained on a highly curated dataset so information that they didn't want in isn't in the pretraining at all. Most other models know the answer and were just taught refusal in post-training.

includenotfound9 days ago
Just give it a web search / web fetch tool
petra9 days ago
But than it's possible to add the knowledge post training, either via RAG, qlora, etc.
Tepix10 days ago
Keep a close eye on abliterated and "heretic" open weight models. They will be outlawed first.
roenxi10 days ago
It is not feasible. They never made much of an inroad against torrents and that is a much easier target than abliterated models. As the linked website shows; the process to abliterate a model can be as simple as

pip install -U heretic-llm && heretic Qwen/Qwen3.5-4B

let alone people just putting the weights up in a torrent. All assuming that someone even tried to ban abliterated models.

Sayrus10 days ago
The torrents you are talking about are outlawed. Whether enforcement is working or not is another issue.
vman819 days ago
Just because something is easily available does not mean that it isn't easily banned. That doesn't make it go away, but it gives a dystopian government a lot of excuses to go after people breaking the law.
api10 days ago
This is the test. If the speech that's easiest to dislike is legal, then we all have free speech.

IMO math is free speech, and outlawing math is censorship.

totetsu9 days ago
At the moment is there any legislature that is seriously pressing regulation to as you say outlaw .. or ban outright AI models that do not have guard rails built in? i know there is a lot of moves about this for things used in critical infrastructure.. but i thought no one is really saying ban these things completely.
Ajedi329 days ago
They picked a good name for fighting that. The optics of trying to outlaw heresy probably aren't great. ;)
kelnos9 days ago
I feel like half the people in the US would zealously defend and support the government if it wanted to outlaw heresy.
anon2919 days ago
You have a fundamental right to multiply matrices.
_0xdd9 days ago
I'll wait for Hexen, thanks.
Bluestein9 days ago
(I must say I thought not the same, but close. That was quite the game.-)
phoronixrly10 days ago
Can the load-bearing gaps that are worth being flagged for pinning down be abliterated out of a model?
chmod77510 days ago
That's the right question to ask. One honest caveat: The interface seam currently forces the pin at the intermediate. Want me to implement or address the other item first?
Bluestein10 days ago
Honest take. Implementing first would break the seams, our work here is done. This is a great place to stop.-

Read the full thread on Hacker News →

Related stories