416 comments
Feels like Anthropic crying do as I say not as I do.
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.
Anthropic’s anger here seems mostly rooted in their annoyance that this exposes they don’t really have core IP that’s not just easily replicated. And that’s clearly a problem for a deeply unprofitable company trying to convince people they’re worth $2 trillion.
Oh that’s very sad.
Meanwhile Anthropic made a product from the work effort of millions of people without compensating them, sell that product on tap and unless I am mistaken do not even have their competitors’ cover of having released any sort of meaningful open weights model.
They have taken from culture (including very specifically their most direct customers’ specific culture — our culture), turned it into a machine to make themselves rich, appear likely to predicate their valuation on permanently removing people from the workforce, then want to dump themselves onto pensions funds and ordinary savers to carry the bag.
It is, I agree, philosophical, because karma is a philosophy as well as a bitch.
Anthropic and OpenAi are spending a $$$$ to "distill" human output into an AI model, then others are spending $$ to distill their AI model into a near-equivalent model.
This is the same reason IP rights exist. On the surface, something like a patent feels ludicrious and even feels morally wrong. Some guy wrote down the recipe for arranging atoms or bits in a particular way, and now I can't!? However, it's designed to solve the same problem, figuring out and describing the process is much more costly than replicating it.
That doesn't mean they won't try, and that also doesn't mean they won't succeed.
I wish people here could at least bother to inform themselves about the IP rights they are so quick to insist are abhorrent, when they seem to not even have a first clue as to what they actually cover.
You could say Anthropic distilled human knowledge and art.
Does anyone know if there are any distillation datasets available? I'd love to see these distributed on BitTorrent. I think it's critical that AI be democratized and not isolated in the hands of a few private companies.
The whole point of AI is to get rid of you so rich people can play with the planet like it's minecraft. Software engineers just think they're special because they're the ones building it like they'll get a pat on the head for being good little servants to the investor class. Or worse, that their portfolios will let them be the gods who rule over ashes.
You ask about distillation but I wonder, is there any training datasets (~ TB-order) available that startup folks in SV use or is it so that everyone has to create their own scraping pipeline ?
While I'd like to agree with this, the fact is that pushing the frontier out has always taken (and folks expect to continue to take) hundreds of millions/billions of dollars. Open source and distilled models can follow on for much cheaper, but it's hard to imagine the frontier ever being "democratized" given the huge sums of money required. It was this realization that forced OpenAI to take tons of private investment in the first place.
The amount of progress that came out of academia and other public sources should not be underestimated either and without all that OpenAI and Anthropic wouldn't even exist.
I think this explains why they are open sourcing broadly. It's not to be nice. It's a strategic play by the Chinese government to help ensure there are many players in this race and not too much power accumulates to American labs (even if American labs benefit in the process)
There is no rule that every Chinese LLM company must open-source their models, and many don't.
(Similar to how U.S. labs have NSA leadership)
But it does not explain publishing techniques like the ones referenced here in this article.
> If China ever gets ahead, they're going closed source and weights immediately.
Anthropic itself admits that Chinese models are merely months behind. Your argument does not make sense, because Chinese labs are contributing massive optimizations like the one this post is about.
Please quote people you appear to be patronizing. China can't do anything about previous released self-hosted Chinese models. If you can show that local Chinese models funnel vast amounts us data home I'm sure you can move a lot of people to your side.
Comments like this also always fail to address why there aren't Western AI companies doing the same thing. Is it because they might get sued into oblivion by Big AI in the US?
It might be better for all of us if you solve that first instead of repeating something the government has been repeating for the last decade or more. It does this, mind you, while sabotaging itself in countless high-tech fields and leaving it all to China for the taking.
Commoditizing one’s compliments is a different strategy.
Or perhaps they've looked at history and concluded that this historically hasn't been where the value is, anyway. It wouldn't be unprecedented - FAANG companies have a long tradition of publishing their algorithms and releasing open weight models. Because they saw the real value as being the training data and in proprietary special-purpose models. For example Google published the transformer architecture and released BERT as an open weight model, but doesn't really even talk in public about the (presumanbly) specialized internal models behind revenue-generating products.
That's giving a lot of credit to Google's organizational ability to productize Google's research...
Local inference will have a boom of cheap, powerful, and available cards at some point (even if it isn’t until 2028/2029). At some point the hyperscalers, and frontier labs, will face the capex problems that everyone talks about, and NVidia, AMD, Apple, and Intel will want to keep selling products.
Powerful, by today’s standard, local inference needs to be accessible to really unlock the “AI” economy long term. It’s just like how the move from mainframes to the PC 40ish years ago unlocked the “computer revolution.”
I realize "going as fast as we can" is not the most popular position atm. But I'm far more interested in what good we can do than 10% apocalypse scenarios. I volunteer with a charity for childhood brain cancer and I do not want to see another 4 year old die. I'm willing to risk anything to stop this.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 5 days ago
- The Verge · 0 points · 3 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 10 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 12 days ago
- I have some questions for Mark Zuckerbergtheverge.comThe Verge · 0 points · 7 days ago