Nvidia CEO Jensen Huang argued that AI model distillation is "competition". U.S. Treasury Secretary Scott Bessent has called it "theft."

72 points•cramer4next•2 days ago•81 comments•

81 comments

captainbland2 days ago
I think it would be totally incoherent to say that copyrighted information is essentially fair game to include in your model but the outputs of your model are privileged against being included in other models.

I'm not sure "competition" is necessarily the right word but I don't think that model creators and their political backers have a leg to stand on when complaining about distillation.

zugi2 days ago
Exactly. AI companies distilled the whole internet into AI weights, now other AI companies are distilling their work to generate new AI weights. Nothing wrong with that, and the world gets open weights AI models.
Alex39172 days ago
> I think it would be totally incoherent to say that copyrighted information is essentially fair game to include in your model but the outputs of your model are privileged against being included in other models.

IANAL, but going from copyrighted texts to an AI model likely constitutes a new original creation, not a slavish reproduction. Whereas going from one AI model to another is more likely to be considered a slavish reproduction reproduction, although arguably AI models are not copyrightable to the extent that the weights are objective facts, like entries in a phone book.

zugi2 days ago
Downloading and redistributing companies' exact AI models would violate copyright. Querying multiple AI models to form your own model with a unique set of weights does not.
captainbland2 days ago
Realistically the companies will organise their data through a network which will involve distributing and storing the copyrighted material without the consent of the copyright holder. This is in and of itself generally ruled to be a breach of copyright irrespective of whether it is fed into a model or not.
thih92 days ago
And this stance isn’t new - AI output is already considered public domain.
SR2Z2 days ago
This is not true and there have not been any court cases which decided it.

It's entirely possible for AI output to be copyrighted if it meets the requirements; prompting an AI takes some skill.

DonsDiscountGas2 days ago
One is allegedly a copyright violation (fair use IMHO but still a grey area legally AIUI), one is a very explicitly a violation of the terms of service you agree to when signing up. And yes I do think that if somebody puts their website behind some barrier where a user has to actively affirm they won't download the contents and use it to train an LLM that should be protected by the same law for the same reason.
sterlind2 days ago
many, many websites have terms of service that forbid scraping. how many of those ended up in the training set, I wonder?
hashmap2 days ago
Are you arguing that training on copyrighted material is fine but distilling against tos is somehow bad? Lmao what a take. Nah, if copyright cant even be enforced then theres no way on earth violating a tos should ever have repercussions other than they ban you. stop carrying water theyre not gonna pay you
FireBeyond2 days ago
So all those authors needed to do was add a TOS page in front of the contents of their books!
mindslight2 days ago
It's so nice of you to think of the poor ignored Terms of Use rather than merely clicking the nag box to proceed like everyone else. Maybe with some more positive attention the Terms might become less depressed, abstain from binge eating, and stop putting on pounds of new paragraphs every year.
ncr1002 days ago
Edit: I realize I can't justify this and it's an irresponsible comment. Please ignore it.

Original:

Smells like a CEO justifying theft. It is reasonable to believe Wong thinks differently about law breaking, given this article.

Something something legal liabilities something something?

captainbland2 days ago
Which one? Seems like there's a lot of it going about.
bennettpompi12 days ago
I think that Jensen's position as the guy selling the proverbial shovels incentivizes him to take a lot of irresponsible positions (namely around safety) but imo he's fundamentally correct here.
verdverm2 days ago
the worst thing I heard him say is that he doesn't think kids should learn their multiplication tables anymore on the Ezra podcast

while I generally agree on his open weight stances, I lost all respect in that moment, everyone of these people are so out of touch

innocent_name2 days ago
>he doesn't think kids should learn their multiplication tables

He said „The majority of the kids shouldn't learn”, not everybody. I listened to that pod and was under the impression that he wants to be portrayed as a "grounded" person and i didn't find his takes to be out of touch.

ks20482 days ago
He comes off bad in this interview. Out of touch is correct - "I had to pump gas once a few years ago ..." and panicked because he didn't know his address or zip code.
dabinat2 days ago
Personally I think multiplication tables aren’t that useful because they’re just rote memorization. I think it’s better to teach kids to memorize a few such as x * 2, x * 5, x * 10 and teach them how to extrapolate from there. Extrapolation is a more useful skill than memorization.
bobajeff2 days ago
As sometime who's memorized but also forgotten their timetables while still managing to go through Algebra I. I don't understand the importance for memorizing times tables (or any other tables for that matter).
stevenwoo2 days ago
AFAIK he was outright lying when he said he forgot his zip code. It hasn’t been possible to buy gas in Bay Area for decades without inputting your zip code for credit card purchases. That smelled of some bad speechwriting.
Eliah_Lakhin2 days ago
Now when the big LLM companies create their products by accumulating a notable portion of copyrighted creative works (including computer programs) from Internet, it does not count as copyright infringement or competition (or "theft"). It is considered as "fair use". So, why training LLMs on other LLMs is not fair use too?

> If you don’t like that, if you don’t like people to use your products, all you [have to do is] know your customers, and disable the service

Considering the current situation, I understand Mr. Huang point, but that's not how the copyright framework assumed to be working from the beginning. It should protect both small actors (authors) and the big companies from unrestricted use of creative works. Now this mechanism seems to be practically dysfunctional.

And a big portion of this lies on shoulders of proponents of permissive OSS, who defend an idea of (almost) unrestricted use of their source code texts for many years, and long before mass LLM scrapping became a thing.

chrsw2 days ago
Would a 100% compatible and open source CUDA stack be “competition” too?
innocent_name2 days ago
Yes, see ROCm.
HarHarVeryFunny2 days ago
I wonder how the "Chinese AI is all just distillation" folk are going to cope now that the Chinese are all aboard the post-train via agents in custom training environments wagon?

Here's Xiaomi's discussion of this, plus their open-sourcing of 7000+ RL training environments.

https://www.alphaxiv.org/abs/2609.mimo-scaling-reinforcement...

https://mimo.mi.com/docs/en-US/news/latest/v2-6

Who needs a few of someone else's "vacation postcards" of their post-training experience when your agents can go on vacation themselves!

Read the full thread on Hacker News →

Related stories