379 points•allisdust•about 12 hours ago•416 comments•

416 comments

cmiles8about 12 hours ago
Why is it a problem that the Chinese labs are just distilling down Anthropic’s models? Aren’t Anthropic’s models not just distilling down other people’s work?

Feels like Anthropic crying do as I say not as I do.

jedbergabout 12 hours ago
What Anthropic is doing requires way more resources than what the Chinese labs are doing. So their complaint is that they do 95% of the work and the Chinese labs do the last 5% and call it their own.

An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.

JackFrabout 12 hours ago
But the analogy still holds.

The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.

cmiles8about 12 hours ago
I get that angle but it’s a weak argument as Anthropic is doing the same to others. Also while there’s certainly a lot of computing power needed to do what Anthropic does, it’s increasingly clear there isn’t much secret sauce involved. Everyone knows how do to the core work it’s just a question of who wants to burn billions on compute to do it.

Anthropic’s anger here seems mostly rooted in their annoyance that this exposes they don’t really have core IP that’s not just easily replicated. And that’s clearly a problem for a deeply unprofitable company trying to convince people they’re worth $2 trillion.

dofmabout 11 hours ago
> What Anthropic is doing requires way more resources than what the Chinese labs are doing.

Oh that’s very sad.

Meanwhile Anthropic made a product from the work effort of millions of people without compensating them, sell that product on tap and unless I am mistaken do not even have their competitors’ cover of having released any sort of meaningful open weights model.

They have taken from culture (including very specifically their most direct customers’ specific culture — our culture), turned it into a machine to make themselves rich, appear likely to predicate their valuation on permanently removing people from the workforce, then want to dump themselves onto pensions funds and ordinary savers to carry the bag.

It is, I agree, philosophical, because karma is a philosophy as well as a bitch.

bushbabaabout 11 hours ago
And the communal work of humanity is orders of magnitude more work than what anthropic pays for their scraping of content. I got no check from them for my contributions
faangguyindiaabout 12 hours ago
Isn't it better for planet? By not doing the wasteful transformation work again
seizethecheeseabout 11 hours ago
The conversation here is mostly moral and ethical but the problem here seems to be financial.

Anthropic and OpenAi are spending a $$$$ to "distill" human output into an AI model, then others are spending $$ to distill their AI model into a near-equivalent model.

This is the same reason IP rights exist. On the surface, something like a patent feels ludicrious and even feels morally wrong. Some guy wrote down the recipe for arranging atoms or bits in a particular way, and now I can't!? However, it's designed to solve the same problem, figuring out and describing the process is much more costly than replicating it.

ASalazarMXabout 11 hours ago
AI training, if viewed through the capitalist mindset, is plain theft. Anthropic can't morally defend copying someone else's IP, but denouncing others copying Anthropic's stolen IP.

That doesn't mean they won't try, and that also doesn't mean they won't succeed.

AlexandrBabout 11 hours ago
Anthropic and OpenAI are very happy to ignore the IP rights of others, so I'm not sure how they can ask for any kind of IP protection themselves. Live by the sword, die by the sword.
failbufferabout 11 hours ago
Capitalism: moral rights exist when they give us a moat.
freejazzabout 11 hours ago
Model weights wouldn't be covered in a patent. You could patent a method of creating weights in a model, but you couldn't patent the weights themselves.

I wish people here could at least bother to inform themselves about the IP rights they are so quick to insist are abhorrent, when they seem to not even have a first clue as to what they actually cover.

jrfloabout 11 hours ago
Because cost of original training >> cost of distilling. It's the same thing that happens with Chinese knockoffs of physical products - it takes a lot of money and R&D time to design a new product, but it's basically free to buy the product, reverse engineer it, and resell it. All the data they originally trained on was available for free on the internet. If the original work was so valuable, it shouldn't be up on the internet for free in the first place imo.
ASalazarMXabout 11 hours ago
Caveat: cost of creating human knowledge/art >>>>>>>>>> cost of original training >> cost of distilling

You could say Anthropic distilled human knowledge and art.

AlexandrBabout 11 hours ago
It's "free" as in beer, not free from copyright. LLMs are free from copyright on the other hand. So which is more "free"?
wonnageabout 11 hours ago
Sounds like Anthropic should close up shop then, those chumps are offering their product on the internet for any random loser to distill
bionhowardabout 11 hours ago
“Distilling” is a funny way to say “learning from”
slowinabout 12 hours ago
I'm also grateful to the Chinese labs for providing workarounds for the walled gardens that the US based AI companies are attempting to create.

Does anyone know if there are any distillation datasets available? I'd love to see these distributed on BitTorrent. I think it's critical that AI be democratized and not isolated in the hands of a few private companies.

10xDevabout 12 hours ago
An authoritarian regime is not your friend and will pullback the moment their own models become highly capable.
pksebbenabout 12 hours ago
Oh no, they might stop doing the thing that benefits me and that they were never required to do in the first place.
computerexabout 11 hours ago
As opposed to what? The US? You think the US is any different? Literally our pedo president publicly admits to insider trading. You think the US government gives a rat's ass about the American people?
slowinabout 12 hours ago
There's no "pulling back" things that have already been open sourced.
horsawlarwayabout 12 hours ago
Yes, we already discussed the US.
voiceofchoiceabout 11 hours ago
Every workplace is an authoritarian regime where workers don't have a say and grovel with "please don't replace me" like that isn't the entire point.

The whole point of AI is to get rid of you so rich people can play with the planet like it's minecraft. Software engineers just think they're special because they're the ones building it like they'll get a pat on the head for being good little servants to the investor class. Or worse, that their portfolios will let them be the gods who rule over ashes.

ducktectiveabout 12 hours ago
>distillation datasets

You ask about distillation but I wonder, is there any training datasets (~ TB-order) available that startup folks in SV use or is it so that everyone has to create their own scraping pipeline ?

forshaperabout 12 hours ago
There are several? And there exists companies whose entire business is just providing them? iirc
frabcusabout 12 hours ago
It's particularly important we all scan and destroy our own unique books!
atherton94027about 12 hours ago
Given the amount of people complaining about crawlers in the past 2 years, I think it's the latter
derpyzzaabout 11 hours ago
there's https://pirateface.co/ which is like huggingface but distributed via torrents
joe_the_userabout 11 hours ago
The Chinese models are to an extent that distillation data.
hn_throwaway_99about 11 hours ago
> I think it's critical that AI be democratized and not isolated in the hands of a few private companies.

While I'd like to agree with this, the fact is that pushing the frontier out has always taken (and folks expect to continue to take) hundreds of millions/billions of dollars. Open source and distilled models can follow on for much cheaper, but it's hard to imagine the frontier ever being "democratized" given the huge sums of money required. It was this realization that forced OpenAI to take tons of private investment in the first place.

jacquesmabout 10 hours ago
That's ok. They took a few trillion worth of content and burned a couple of hundred billion of their own money. I really don't see the problem, if they didn't think about this long and hard before they went down that part it should not be on the rest of us to bail them out. The 'frontier' is less important than the democratization process.

The amount of progress that came out of academia and other public sources should not be underestimated either and without all that OpenAI and Anthropic wouldn't even exist.

bwest87about 12 hours ago
The best explanation is that it's a goal of the CCP to generally commodotize LLMs, because LLMs will ultimately be a compliment to manufacturing (which China dominates), and you always want to "commodotize your compliments".

I think this explains why they are open sourcing broadly. It's not to be nice. It's a strategic play by the Chinese government to help ensure there are many players in this race and not too much power accumulates to American labs (even if American labs benefit in the process)

yorwbaabout 10 hours ago
It's a bad explanation because it assumes decision making about LLM releases is centralized in the CCP even though every AI lab has a different strategy. Some publish LLM research with small models but seem to be staying out of the race to the frontier (Weibo), some train large models but publish few research details and no weights (Bytedance, iFlyTek), some decide on a case-by-case basis what they publish or not (Alibaba, Baidu), ...

There is no rule that every Chinese LLM company must open-source their models, and many don't.

singularity2001about 8 hours ago
You are aware of the rule that any company has to have a CCP representative in the leadership?

(Similar to how U.S. labs have NSA leadership)

isignalabout 2 hours ago
One explanation could be that they know Western consumers will never use the llms directly, so they can open weight the models and collect licensing fees from hosting providers in US who run it. Moonshot seems to have such terms in their open weights releases.

But it does not explain publishing techniques like the ones referenced here in this article.

jrfloabout 11 hours ago
Totally agreed. People are so ready to praise China for their free models, but they aren't doing it because they believe in free open-source software. If China ever gets ahead, they're going closed source and weights immediately.
computerexabout 9 hours ago
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

> If China ever gets ahead, they're going closed source and weights immediately.

Anthropic itself admits that Chinese models are merely months behind. Your argument does not make sense, because Chinese labs are contributing massive optimizations like the one this post is about.

sillyflukeabout 11 hours ago
>People are so ready to praise China

Please quote people you appear to be patronizing. China can't do anything about previous released self-hosted Chinese models. If you can show that local Chinese models funnel vast amounts us data home I'm sure you can move a lot of people to your side.

Comments like this also always fail to address why there aren't Western AI companies doing the same thing. Is it because they might get sued into oblivion by Big AI in the US?

It might be better for all of us if you solve that first instead of repeating something the government has been repeating for the last decade or more. It does this, mind you, while sabotaging itself in countless high-tech fields and leaving it all to China for the taking.

layer8about 11 hours ago
*complement

Commoditizing one’s compliments is a different strategy.

ourafabout 5 hours ago
A very effective strategy among Middle Management, i might add
bunderbunderabout 11 hours ago
It's also possible that China has decided that this is ultimately going to be a race to the bottom, anyway, and values the soft power more highly than potential monetary profits.

Or perhaps they've looked at history and concluded that this historically hasn't been where the value is, anyway. It wouldn't be unprecedented - FAANG companies have a long tradition of publishing their algorithms and releasing open weight models. Because they saw the real value as being the training data and in proprietary special-purpose models. For example Google published the transformer architecture and released BERT as an open weight model, but doesn't really even talk in public about the (presumanbly) specialized internal models behind revenue-generating products.

ethbr1about 10 hours ago
> For example Google published the transformer architecture and released BERT as an open weight model, but doesn't really even talk in public about the (presumanbly) specialized internal models behind revenue-generating products.

That's giving a lot of credit to Google's organizational ability to productize Google's research...

reedf1about 12 hours ago
I've been running Qwen 3.8 27b (an opus 4.6 tier model), locally on a 5090 for just over two weeks @ 170 tokens/s. That's a frontier model from 9 months ago running on consumer hardware. Who knows where distillation and pruning gets us in another year.
bix6about 12 hours ago
$9k for a 5090 now? Sheesh.
rubyn00bieabout 11 hours ago
In all fairness there are probably a lot of folks who picked one up for around MSRP (even if one of the board partner cards with an MSRP 10-15% over the FE).

Local inference will have a boom of cheap, powerful, and available cards at some point (even if it isn’t until 2028/2029). At some point the hyperscalers, and frontier labs, will face the capex problems that everyone talks about, and NVidia, AMD, Apple, and Intel will want to keep selling products.

Powerful, by today’s standard, local inference needs to be accessible to really unlock the “AI” economy long term. It’s just like how the move from mainframes to the PC 40ish years ago unlocked the “computer revolution.”

bitexploderabout 11 hours ago
Well, I have a $750 card that runs at about 50-60% of that token rate :)
off_with_their_about 12 hours ago
$9k is a small price to pay to experience the rapturous glory of AGI. I'd easily pay up to 3 times that to comfortably run the superintelligent models released in this post RSI world.
zdragnarabout 11 hours ago
Weird, I kinda gave up on 3.8 as anything other than a planner. I had it try to write some basic unit tests for an admittedly complex bit of code and it ran out of context thinking about the problem and exploring random parts of the code base repeatedly before it even wrote a single line. Toning down the thinking helped some, but then it wasn't much better than qwen coder.
oidarabout 12 hours ago
What are you thoughts on it's performance compared to 4.6?
reedf1about 12 hours ago
Indistinguishable or very mildly better. But it's considerably faster. Some portion of that is also probably down to improvements in model harnesses, I've been using opencode.
newyankeeabout 11 hours ago
Do you think this trend can continue ? An Opus5.5 equivalent on a slightly bigger local hardware in under a year ?
an0malousabout 11 hours ago
I’m not an AI researcher, but it seems like there’s a ton of waste having a universal model that knows everything when any individuals use case requires like generously 10% of what’s stored in the model. Does it even need to have memorized knowledge stored in the model or could it just look up info and docs like humans do? If all you need is the language and intelligence, I think Opus5.5 equivalent intelligence will run on an iPhone within 5 years.
redanddeadabout 12 hours ago
Well how’s it been so far
listlessabout 12 hours ago
I'm beyond thankful that Chinese AI models are so good. I desperately want us to cure the myriad of maladies that humans suffer needlessly with on a daily basis. We're going to need more powerful models than we have now if we're gonna do that and the Chinese are providing the competition needed to push this thing as fast as we can.

I realize "going as fast as we can" is not the most popular position atm. But I'm far more interested in what good we can do than 10% apocalypse scenarios. I volunteer with a charity for childhood brain cancer and I do not want to see another 4 year old die. I'm willing to risk anything to stop this.

networkedabout 11 hours ago
Do you mean that you don't believe in the 10% apocalypse scenarios or that you think they're an acceptable risk? Only the latter is really "risk anything".
listlessabout 1 hour ago
I would say I’m willing to accept the risk. But I do believe the risk is overblown. It’s always overblown.
dgellowabout 10 hours ago
I’m pretty sure they meant they believe in the 10% risk of everybody dying but are ok for all of us to take that risk without our consent because they saw a 4year old tragically die. Which sounds completely unhinged, to say the least
idbnstraabout 11 hours ago
i don't know much about medicine so i'm curious about how you're using AI, and how medicine in general is using AI
listlessabout 1 hour ago
I don’t actually know. But it’s logical to me that a system designed to find answers in context is well suited for these healthcare problems. We know the answer is somewhere in all the data we have. It’s just too difficult and complex for us to get through quickly.
n1b0mabout 11 hours ago
But you’re ok with AI being used to kill school children in Iran?
sixoabout 11 hours ago
This really is a place where the guns-don't-kill-people argument applies, even moreso than guns themselves. The U.S. government massacred school children in Iran. Why does it matter how they targetted them?
mikeg8about 11 hours ago
Most people aren’t okay with that, but the blame lies in the people who deployed the AI tool in that situation, not the makers of said tool.

Read the full thread on Hacker News →

Related stories