Newly unsealed court filings show Microsoft privately called OpenAI's data practices "theft" while both companies scraped paywalled Times content, built datasets from it, and warned internally it would gut…

955 points•pluc•13 days ago•841 comments•

841 comments

haritha-j13 days ago
I just don't understand people saying "but a human learning from a book isn't illegal".

How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."

And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.

coffeefirst13 days ago
Because it’s a bad faith argument that presupposes integrating someone else’s work into your algorithm is equivalent to me reading a book.

Your rule would be the right way to do all this. You could even have a mechanical royalty that applies by default where you can train on anything that hasn’t set rules and a preset rate.

ryandrake12 days ago
It's the same bad faith argument as:

"It's perfectly OK for a police officer to observe a street corner, see crime happening, and go take action, therefore, building a complete, panopticon surveillance system that watches all street corners simultaneously and deploys police to take action, is perfectly OK, too, since that is exactly the same thing."

4d4m11 days ago
Correct take according to law. A mechanical style royalty would be requisite
kenjackson12 days ago
"How do people not understand that some laws only make sense at a certain scale? "

Honestly, this is because its something that is basically never discussed or reasoned about. The closest I can think of is "personal use vs commercial use". But I'd love to see more about how to reason about how laws change at scale.

devsda12 days ago
Laws do consider scale and we can see that in everyday law.

4 friends walk together, it's just normal. 400 "friends" walking can and will be treated differently.

Moving around with a couple of bills is treated differently from carrying huge bundles of cash.

Those laws or exceptions were probably added later as a reaction to abuse of existing laws.

The problem with AI/scraping is that we can't afford to be reactionary because it may be too late by the time we realize what has happened and change the law.

It will be too late because governments and judiciary have been largely about maintaining the status quo and minimizing disruption when it comes to big tech related cases even when they have been found guilty of wrongdoings. We already see the too big to fail vibes with AI.

morkalork12 days ago
Murder, terrorism and genocide.

Simple possession of drugs vs possession with intent to distribute. In jurisdictions that make difference base on quantity (so, scale) alone.

I'm sure there's other examples too.

leobg13 days ago
Because copyright is a tradeoff. Society wants information to be free. But authors won’t publish if anyone can republish their work for free. Hence the distinction: The ideas you write are not protected - only your expression is.

In a sense, AI changes nothing. Society profits from having AI just as it profits from having people learning from others. In both cases, those who stand on the shoulders of others still make money for themselves. But the economy as a file is richer, too, because people can choose to buy something better now that wasn’t available before.

vharuck12 days ago
Allowing AI companies to get away with this because we're "better off"¹ is like eating our seeds. Sure, we built something cool with the massive corpus available, but how will it affect future decisions to develop or share creative work?

If nothing else, a sense of justice tells me that if somebody's work directly helps create a profitable tool, that person should share some of the profit. The size of the share can be negotiated, but the AI companies didn't even reach out before the law suits. And even then, only to major sources of content (some of whom don't have the copyright for their content, just a limited license for distribution on a website and all the nitty gritty involved in that).

1: In a sense, a world with these models has more capabilities and is therefore better. But this is the real world with real people, who are emotional and competitive. So let's see how it actually plays out.

dhx12 days ago
> authors won’t publish if anyone can republish their work for free

There's plenty (even a majority?) of authors that publish and will continue to publish without any expectation of direct remuneration. Open source software developers and companies hiring such developers. Not-for-profit organisations increasing awareness of a cause. Private companies wanting to reach an audience for marketing reasons.[1] Government organisations. Researchers funded by government grants. Universities publishing books or coursework openly (they're in the business of selling their stamps on degrees, not selling books).

[1] Even includes the likes of Warner Music with CC-BY music videos on YouTube for some artists, seemingly for marketing reasons to try and build the name and following of a particular artist.

make312 days ago
>Society profits from having AI

Not all of society at all. Let's discuss this when AI is better integrated and a large fraction of people are laid off in 5 years.

moscoe12 days ago
I really don’t understand why anyone would think they are entitled to any part of a derivative of a published piece. Why publish if not to help advance humanity, just like all who came before you and contributed to your success/ability? This is quite literally the bedrock of the progress of human civilization. As long as they aren’t straight up reproducing a direct copy of the original work.

If you want to keep it to yourself so only you can benefit from it, then keep it private. Otherwise, why don’t you get to work on the next big idea.

rrr_oh_man12 days ago
> As long as they aren’t straight up reproducing a direct copy of the original work

I think that part is debatable

gigel8212 days ago
That's eerily similar to the argument used by mass surveillance systems like Flock around expectations of privacy in public.

Yes, it's fine if the little old lady around the corner writes down the color or plates of cars driving through the road a few hours a week, but no, it's totally not fine for an all-seeing, all-powerful entity to collect all license plates, and photos of drivers and passengers, with exact metadata to automatically process and sell that data for profit to anyone who would pay.

4728284713 days ago
“Information wants to be free“.

It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.

My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.

proc013 days ago
"I just wished the collected data was public. "

That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.

Gareth32113 days ago
I agree. The double standard is the problem. People have been imprisoned for IP theft, but when these companies commit IP theft on the grandest scale ever imaginable, they're rewarded with trillion dollar IPOs. Either IP isn't protected, or it is. Legislators need to pick a lane. Right now it appears that poor people go to prison, and rich people get rewarded.
blfr13 days ago
Other companies have no qualms about distilling the first. Let's hop on gear and get the market to deliver a distilled Fable that runs on a smartwatch. Sooner is better.
CJefferson12 days ago
Anthropic getting angry other AIs are trained on their AIs output is, to me, one of the stupidest things I’ve read in a while.
jrflo13 days ago
I think the difference is that the companies are dumping billions of dollars into transforming that data into something useful, so they would like a return on their profits. Opening up the models for free is not a good business model if you want to make money.
DownGoat13 days ago
On the flip side, output from an LLM is not copyrighted.
polytely13 days ago
I sort of agree, and i think strengtening IP Law is probably not great. But I do think it's very fucked that building generative ai is only possible by taking the works of countless artists and craftspeople and then the model produced from that data immediately gets deployed to destroy the careers of the people whose, work was vital to it being created, without compensation for them, while making a few evil nerds richer than god. I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.
juiceland13 days ago
> I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.

Now we’re getting somewhere. Let’s start with redistributing the profits from AI companies and then move on to all profits from all companies because the logic is the same.

fwlr13 days ago
Well the future we seem to be getting is “information wants to be free for the first ten thousand tokens, then $1 per million tokens after”.
TeMPOraL12 days ago
Your point notwithstanding, that's still a bargain.
alentred13 days ago
> Sharing information is an act of love

Most AI companies are not sharing it, though. They appropriated it and resell it.

cush12 days ago
Yikes. That’s some deep entitlement.

Unfortunately in the real world there’s this thing called money, and we exchange it for goods and services. The reason information isn’t free is because it costs time to produce it and people need to be fed.

If you believe that a creator doesn’t need to consent and doesn’t deserve credit or compensation for their work, then you’re likely not someone who has many fundamental needs unmet

These AI companies actively chose not to get consent from creators and earn billions from their content with no compensation.

r3trohack3r12 days ago
IIUC, the question at hand is: does training require a special, separate, license or can you legally acquire a work and then use it for training?

I.E. Anthropic can not pirate a bunch of books and then use those for training, but it can legally purchase the same books and then use those purchased books for training.

armchairhacker12 days ago
In the past, but today fewer people are getting paid less this way.

Would you rather resurrect IP law, or find some new way to pay creators, then finish killing it?

sajithdilshan13 days ago
If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
y-curious13 days ago
Yeah gulags and other forced work camps also come to mind. But I guess this is a larger scale in terms of man hours
bcjdjsndon13 days ago
But it's copying...how is it theft? Your labour WASNT stolen was it?
mitxela13 days ago
Never ended, just changed in form.
not_a_bot_4sho12 days ago
You're right that it never ended.

But it didn't change form much. Still around 50 million people enslaved nowadays.

midtake12 days ago
Are you comparing modern workplace aches and gripes to literal 1800s slavery?
bcjdjsndon13 days ago
No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters
xxs13 days ago
Slavery is a weird one. It has been there for longer than any written history exists. In ancient times (Greece, Rome), slaves didn't have rights at all. A horrific injustice but it'd be not be a theft. Then you get the serfdom in the middle ages. Up to recent times humans have been brutally exploited.

The copy part was a recognized right, then taken away.

juvvel13 days ago
I wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everything is getting privatized -- housing, water, electricity, and now, thinking and knowledge. We are heading towards a world where you have to pay even more excessive fees just for existing and for completing any basic task.
Draiken13 days ago
It's the age old privatize the profits and socialize the losses.

People lose their jobs, the environment is destroyed, our bills skyrocket and all of the gains go to the people who own all the shit...

I honestly cannot believe some people still believe that we'll ever get to a society where nobody has to work and we can live our lives happily ever after. Maybe too many Disney stories?

sleight4212 days ago
For the life of me, I can't understand why your comment is being downvoted or flagged or whatever makes it go gray on HN.

I'm guessing it's people knee-jerking that you're being political?

I can't understand the people who don't see it.

The data centers strain the power grids then electricity costs go up for everyone else. This is de facto a regressive tax because everyone needs electricity and the poor pay proportionately more of their income for the increased cost.

Live in San Francisco? Probably not now unless you're rich because of the skyrocketing cost of living due to Tech and AI money. Another de facto regressive tax, driving away other people.

Environmental damage? The poor are the most impacted and the least able to absorb the costs. Do they have the property or renter's insurance to protect them from these disasters? Another de facto regressive tax.

Need a new phone? Or a computer? Same problem.

Want to dabble in AI? You're not going to get too far on $20/month. It's mostly a wealthy person's game.

Or there are the statistics that the vast majority of successful founders from up upper middle class families or wealthier.

Wealth centralization is what our economic system does. The purpose of a system is what it does. If it wasn't, the system would have been changed.

lofaszvanitt12 days ago
People are like sheep. Plus those who work are preoccupied and are too tired to react to these changes.
TutleCpt13 days ago
The most shocking point is that they have a Microsoft exec who knows what he's talking about.
vintagedave13 days ago
In my experience many execs know what they're talking about.

Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.

Knowledge alone is less often a factor.

Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.

Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.

ekunazanu13 days ago
I think a different variant/opposite of Hanlon's razor applies when it comes to corporate or political decisions: Don't attribute to stupidity when it can be adequately explained by malice or greed.

This sounds rather obvious, but I feel people forget it far too often.

soraminazuki12 days ago
You can see the same dynamic playing out in this very thread too.
Sohcahtoa8212 days ago
Oh, I think MS execs (and all other execs, top-level politicians, and pundits for that matter) know what they're talking about, but they'll tell whatever lies they need to enrich and empower themselves.
someguynamedq11 days ago
Do you think msft market cap is just an accident, or...?

Read the full thread on Hacker News →

Related stories