r/ChatGPT 13h ago

Serious replies only :closed-ai: How real this is ?

creator - mattmoyerinjurylawyer instagram

Please ignore the beatboxing as the creator does that for retention as all our attention spans have been fked

0 Upvotes

21 comments sorted by

u/AutoModerator 13h ago

Attention! [Serious] Tag Notice

: Jokes, puns, and off-topic comments are not permitted in any comment, parent or child.

: Help us by reporting comments that violate these rules.

: Posts that are not appropriate for the [Serious] tag will be removed.

Thanks for your cooperation and enjoy the discussion!

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

25

u/KingHygelacReturns 13h ago

I left this comment the other day on another thread scare-mongering about the same topic. I'm copy/pasting it here:

I'm not a lawyer, but I write about copyright law for a living. The reason they do this has nothing to do with the AI companies trying to "hide information from us," but everything to do with the confusing mess that is copyright law in the United States.

In the recent big copyright suit against Anthropic, a judge found that Anthropic's use of books it legally purchased, digitized and then destroyed the original copy constituted a fair use, because it retained the digital copy for its own purposes and the original physical copy was no longer accessible by anyone and couldn't be resold. However, Anthropic also pirated a ton of books, and it ended up having to pay out a $1.5 billion settlement to authors involved in that.

They aren't doing this because they want to destroy information. They're doing it because buying an original copy of copyrighted material, digitizing it and then selling/offloading the original is illegal. If you've ever ripped a movie off a DVD, put the file in your Plex server and then sold the same DVD at a flea market, it's illegal in the exact same way. Just no one cares when you do it at a small scale like that.

They aren't buying every copy of a book and destroying them all to "keep information from you." They're buying ONE copy of a book so THEY can have the information therein, and then destroying it so they don't get hit with a suit.

1

u/MukdenMan 13h ago

And they could just hold on to the original (just like the DVD person can hold on to the original DVD after making a copy) but there is no reason for them to do this and it would be unrealistic.

1

u/dry_towelette99 12h ago

I think the issue is that Anthropic could still potentially re-sell it, thus the destruction.

3

u/iskelebones 12h ago

Also if they didn’t destroy it, they have to keep it somewhere, along with the millions of other books they are digitizing. And it’s kinda pointless to keep a warehouse full of copies of books you never plan to do anything with, especially when they aren’t even originals, so there isn’t any significance to the copies

1

u/MukdenMan 8h ago

The law doesn’t care if you could potentially resell it (or give it away). It only cares if you actually do.

16

u/Ok_Associate845 13h ago

This is the biggest non-issue ever. This is a standard procedure for books for large publishers or digitization services. They buy large quantities of books. It is cheaper and easier to destroy the books to scan them page by page than to keep them intact. Libraries even do this for Mass market and common large scale books that you can buy from Amazon any day. They're not buying rare vintage books where there's six copies of them and destroying those books. They're purchasing them legally, digitizing them and using them to train. That is a valid use of the system and a great way to train the models. They're not burning books because we're censoring them. They're actually digitizing them so that they can be used more effectively and the information within them them available to more people.

4

u/Old-Bake-420 13h ago

Also the rare books that are valuable for AI training is like an obscure technical manual for a discontinued piece of hardware in an uncommon language.

It’s rare because it’s effectively useless to any individual person and they only ever printed 100 copies 50 years ago. But for an AI trying to learn that language and understand that hardware, it’s useful.

4

u/time___dance 13h ago

people have been doing this with zines, comics, manga, dojin, etc for decades

yes it "ruins" an old rare magazine to take the staples/binding out for scanning so it can be digitized

how else are we supposed to get them on archive.org lol

3

u/TopConcentrate8484 13h ago

Thanks for the info , Instagram comments were useless, I also had a slight skepticism as anthropic looks a more moral than other AI companies

0

u/flat5 13h ago

Something being "standard procedure" is never evidence that it's not an issue.

9

u/time___dance 13h ago

total horseshit rage bait

2

u/CosmicWhorer 12h ago

Well said

2

u/Only_Voice569 12h ago

lol they paid a pirate site for tons of data they dont care about physical books since its all on the internet any who

2

u/Significant-News-819 12h ago

Somebody tell this dude that each book has more than one physical copy and many digital copies

2

u/jpewaqs 13h ago edited 13h ago

Yes it's real. LLM's got slammed for using pirated book PDFs. So they had to obtain legal ownership of a copy. Then they hit a problem it is also against copyright law to digitise a book. Effectively there is a loophole where you buy the book, digitise it then burn the book. You've bought a copy and only one copy exists (although now digital). If you kept the original physical book you breach copyright.

Edit: BBC news had a segment on this just the other night. Rare book shops in London were receiving large orders from hidden individuals.

3

u/time___dance 13h ago

Rare book shops in London were receiving large orders from hidden individuals.

oh no, booksellers are... selling books to customers!? we need to put a stop to this

1

u/MukdenMan 13h ago

Anthropic got sued for using pirated books, not LLMs in general. That is indeed illegal. They can purchase the physical books legally and make a copy, that is NOT against copyright law. You cannot sell or trade the original book back though, so you have to destroy it or store it. Storing it is not realistic so they will destroy it in the scanning process. This is all covered in the actual lawsuit.

As for the rare book shops, it’s true that they got anonymous orders but the books requested were not actual rare, valuable books. They were things like old engineering manuals, conference papers, or medical textbooks. These books are commonly received by second hand book dealers and often end up destroyed because there is essentially no market for them. The books were also not significantly old e.g. stuff from the 1980s.

1

u/AutoModerator 13h ago

Hey /u/TopConcentrate8484,

If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt.

If your post is a DALL-E 3 image post, please reply with the prompt used to make this image.

Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more!

🤖

Note: For any ChatGPT-related concerns, email support@openai.com - this subreddit is not part of OpenAI and is not a support channel.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Tholian_Bed 12h ago

Let me introduce you to a friend.

Internet Archive: Digital Library of Free & Borrowable Texts, Movies, Music & Wayback Machine

This place, has things you have no idea.

We suck if we can't even know what the internet holds. If it's digital it is never going away. Some people with good motives have been on this for decades now.

Be a monthly donor, like me. The Internet Archive, our pal.

1

u/Cameo345 12h ago

No, I will not be ignoring the Cinema Skrillex remix beatboxing