It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.
Instead of fixating on a remedy that seeks to criminalize AI, maybe focus on the relatively rather achievable goal of redistributing LLM gains. Would that not be the most desirable justice? What is your alternative, and would you foreclose the future in the name of a past that never really existed in the first place?
Let me fix that for you: A world in the the value of the labor of each human approaches zero.
Humans with zero economic value can still vote. They can still mass. They will still have needs. Really the script here writes itself. The historical precedents bound the problem rather well. As always, radical social change will not occur until a wide swath of the population is aligned. In this case, due to their broad economic devaluation.
It would be easier if today's knowledge workers stopped deluding themselves that their standards of living will maintain. Your acknowledgement of the necessary predicate to change is, in that respect, progress in itself. There is little reason why we cannot accelerate the timing of broad consensus if more people so readily came to that conclusion - and resisted the temptation to then find the answer instead in nostalgia about the past.
It is however, achievable. Certainly more so than engaging in the fantasy that we can criminalize LLMs out of existence. And it is likely more desirable than such an effort anyway.
Nothing about banning technology. How about enforcing DMCA and then applying a penalty for the knowing theft rather than negotiating a license?
We all experience the substance which is trickling down.
Far more likely is that wealth and power will become even more concentrated into the hands of the few and the rest of humanity will become effective slaves.
We are actually going to get the worst of both worlds.
A) people working B) people working more for less C) soulless jobs widespread
What an idiot you are.
Redistribution? Hahaha.
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
really not the same entities here
https://www.napster.com/blog/napster-heads-to-microsoft-buil...
Free means the same as worthless, which inherently isn't true - since information takes time to consume in some form, and your time isn't worthless. Therefore even if you could listen to all songs theoretically for free, you would need to spend an inordinate time doing that.
When I was a kid, getting a CD from your favourite band was a major expense, getting a video game even more so. But it formed a sort of emotional attachment (and not even just for me), my friends talked about how 'band X''s new album was amazing or a stinker. Since there were multiple bands making similar kinds of music, choosing to be a fan of one but not the other carried real monetary weight.
Nowadays you just fish out a song you think you would like out of the endless sea of Spotify, no different from prompting an LLM. No, Spotify didn't make me enjoy music more.
Same applies for Steam & videogames.
Therefore I think the ritualistic act of paying money to get access to something does have a purpose. It inherently establishes the value of information to you, makes it an investment that you need to recoup by using it. I'm sure most musicians would trade a million fans who might check them out if they're in town, to ones who think their music changed their perspective in life.
Also the process of creating a song that vaguely appeals to millions is different from making one that speaks to a thousand.
This is a fundamental issue of modern capitalistic society, similar to the Marxist idea of 'alienation' - once something is cheap to get, you don't appreciate the effort that went into making it. And if your customers don't care about the thing they get, producers won't make an effor to make it good either.
And once nobody cares, people even forget what a quality product is like.
One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
Did you miss the "book burning" hysteria from a couple weeks ago? These companies have been trying to digitize copyrighted materials legally, in which copyright law demands destruction of the original, and people shit on them even harder.
It's clearly not a problem for these companies to buy the books they need for training, and they have been doing that in crazy high volumes. Lots of good training materials simply cannot be legally purchased though, and should those parts of human knowledge just be ignored?
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
For instance a image/video generating model.
So whose viewpoint is right here? Is downloading theft or not? These arguments always boil down to "it's fine when I do it, but wrong when a company does."
The problem is that it is enforced exactly the opposite. People have been hit with fines and jail time for pirating and seeding, without even doing so for commercial gain. But when massive tech companies pirate training data for their AI and build a product from that that, nobody goes to jail. Where is the sense in that?
It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.
We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.
EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.
If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
Any argument that writers and artists lose from these existing, would remain unchanged.
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.
The 'sell it back to us' argument falls short in my view.
Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones.
The comment here seems incredibly pessimistic and quite dramatical.
Awesome, can I make my own competitive LLM, just like I can make my own open source software?
> and in some time useful models will ship preinstalled on all mobile phones.
Considering hardware prices, that "some time" is doing super heavy lifting. It could be 10-15+ years before that happens and the local LLM is actually useful. Most people don't see hardware prices declining from current prices until at least 2030, likely much longer.
Yes, it could take some time to arrive on phones. It is questionable if it will ever make sense compared to using a paid hosted provider.
But what are 10 years in the grand scheme of things?
Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?
There will be new jobs coming from this.
The bountiful abundance of intelligence is truly the best thing that has happened this decade.
It's asinine that you think the sell it back to us argument falls short.
Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.
From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.
We have social conventions regulating this.
Suddenly, mega corporations were allowed to digest(sometimes by illegally pirating “data”, and sometimes by achieving their training corpus and subsequently destroying the copies) and digest this information in a novel way, without any discussion or law making.
You may argue it’s beneficial (it very well could be, I use ChatGPT and Claude all the time), but let’s not pretend it’s the same as learning from your history teacher…
Pirating is fine.
Subsequently destroying is bad but it is the result of screeching from people lacking foresight who support the idea of training LLM on book copies is wrong. So now corps found "legal" way to do it by destroying it. Again, idiotic social conventions.
Without discussing? Law making? What are you, German? There's a reason EU is shit while USA is center of the progress of the world: laws follow innovation; not the other way around.
The average human doesn't do much on his own, and definitely doesn't uproot society or risk siphoning/leeching wealth from every person on this planet.
Average human on his own...why draw the line there? It doesn't matter much what one human does...but what many/collective/society do and society has been "ripping off", "uprooting", "leeching (read: creating)" wealth since dawn of time. It is called PROGRESS.
There is no such thing as PROGRESS for progress' sake.
And FYI, agriculture is a wonderful invention.
Yet for about 5000 years after its introduction the average human had worse nutrition than the average hunter gatherer, which led to such things as height decreases for those 5000 years.
Industrial agriculture is another wonderful invention. Yet 150 years later we're not sure it's sustainable and it's likely many of its aspects aren't, which will raise some sticky issues soon ("which billion people do we decide to let starve since we can't make enough food for everyone after most of our soil eroded?"). Repeat this for industrial textile production, mining, etc.
I won't even go into climate change.
And again, scale matters. Most individuals can only control what they do, and what they do generally doesn't impact much. But companies can impact a whole lot.
> "ripping off", "uprooting", "leeching (read: creating)" wealth
Let's not be 100% cynical here. A lot of what humanity has achieved has been genuine wealth creation and distribution/re-distribution. I would say more wealth has been created than leeched off.
* * *
And before you think I'm some starry eyed teen, I'll play the game. At the end of the day, me and mine have to outrun you in the face of PROGRESS.
May the odds be ever in your favor.
The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owned anything they invented or created. [0] It was a long time ago and I remember feeling sad watching the story. In the Soviet Union, a group of ~15 people, Politburo, controlled everything including any thought written to paper.
It is this one line, Article 1 Section 8 Clause 8, that separates the United States from the disaster that was the Soviet Union:
> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
I don't think it is far fetched to call ignoring and disregarding the Copyright Clause a communist revolution, violent or not. That is the one thing the communists -- there have been many over the years inside the United States -- would change to make the United States a communist country.
[0] https://en.wikipedia.org/wiki/Tetris#Spread_beyond_the_Sovie...
It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.
This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you hate bigcos and capitalism.
Learning isn't stealing, regardless of whether it's done by a human or a machine. By your logic someone reading and memorizing all somebody's life work is appropriation; completely inane.
The vast a majority of text that is being claimed to have been "stolen" was never for sale. Reddit posts, deviant art images, personal websites, etc.
What if it was for free, like Wikipedia?
> Crimes this large are crimes against humanity.
jfc no, sit down.
Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch.
At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress at the edge of our understanding; there's just too much shit for one person to learn "manually" (wait I'm not advocating for low-effort slop, chill)
It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit.
Example:
Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads.That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website?
Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn't "fit" the community?
God hell no, if you don't know how much bullshit AI could eliminate for the silent majority then you were probably part of that bullshit.
(that's a general "you" for whomever was fine with the status quo and not a personal insult @ anybody)
You sound like one of those mentally unstable vegans who loves their avocado. How cute the animal has to be for them to care...?
In the long term though I think models have no moat, so the cost will fall to the cost of compute and storage. Which is why they’re pushing AI safety panics: regulatory capture to outlaw open models and outlaw competition.
And yeah, EA is neither effective nor altruistic. It’s a cult, part of the “Rationalist” and adjacent cluster of tech cults. They’re to tech what Scientology is to Hollywood I guess.
Yeah the introduction of copyright was truly criminal.
> So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
Oh wait ...
lol at this edgy 5th grade statement. So ridiculous.
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.
Normally this isn't the case for any technology except for the time it first comes around. AI is only different to use for two reasons. First, it is in our time. Second, it seems to be faster than any of the options before, so the shock is harder.
But in general, this is a website of people writing code. How many on here study how a person solves a problem and then trains the ultimate chimpanzee to do (at least part of) their job? Is building computer programs that automate what others did manually theft?
Consider the origin of the word "computer" itself, a mass theft of jobs that would have employed the whole world many many times over.
The copy part was a recognized right, then taken away.
Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
https://arxiv.org/html/2408.02487v3
I wonder how would Microsoft react if someone would synthesize a code solution based on Windows source code.
It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow.
If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.
Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it"
Probably indicative of America's wider downfall that they've all become so self interested
If it has no soul to save and no body to torture it deserves no rights.
It is really insane to compare individuals copying data to big corporations parasiting on the Internet.
If this was all open, I’d maybe half agree.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
Might be a short one though if all goes to plan. Just another form of gatekeeping the worlds information and with new gatekeepers replacing the old ones.
> At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
Look at who owns that site, no surprise here.
Publicly. Accessible.
Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic.
Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.
If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.
If I build a complex custom bike and you start copying individual features from it on your custom bikes, surely you are committing theft of some sort, but whether it’s punishable depends on whether I’ve decided to go full corporate and protect my designs with patents and trademarks. You’ll be hard pressed to patent or trademark anything if I have published and documented prior art.
I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.
Honest question. There is a line in the sand somewhere apparently.
> "We just need to launder it through a fine-tuned codex." [0]
[0] https://cybernews.com/news/midjourney-ai-images-art-lawsuit-...
AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it.
At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.
At least it’s consistent, is what I’m saying.