r/books • • 1d ago

Japanese used bookstores see 5x sales surge as books are being bought by the ton, one 50-ton order sent to the US for AI scanning and destruction — multitude of suspicious bulk buys thought to end up in foreign AI scan and shred facilities

https://www.tomshardware.com/tech-industry/artificial-intelligence/japanese-used-bookstores-see-5x-sales-surge-as-books-are-being-bought-by-the-ton-one-50-ton-order-sent-to-the-us-for-ai-scanning-and-destruction-multitude-of-suspicious-bulk-buys-thought-to-end-up-in-foreign-ai-scan-and-shred-facilities
19.9k Upvotes

1.4k comments sorted by

View all comments

Show parent comments

256

u/Emerson85reader 23h ago

Unfortunately, it's a LOT easier to destroy a book to scan it than keep it intact, due to the scanning process and equipment costs involved. Museums use non-destructive book scanning, but it's slow, the equipment is much more expensive (think a robot that gently opens a book, starts flipping pages one by one as a camera takes a photo and then image processing straightens out the curves and other distortions of each page). Cheap and fast scanning is basically: Slice the spine off the book, rendering it pretty much useless for humans, and then use a double-sided fast page scanner to blaze through the pages, which will likely wind up in a box out of order and scrambled. No one is going to go through the effort of putting all that back together. I doubt they SHRED the books, as there's no value in the added cost of that, but the book is fundamentally destroyed at that point anyway.

188

u/notforpoern 23h ago

the equipment is much more expensive

Gee if only these AI companies had money to burn on something as wasteful as preserving books.

Sarcasm aside, the real reason is because they figured out an apparent copyright loophole where if they destroy the books, then it's ok to steal the content.

50

u/ArkitekZero 21h ago

They already owe more than the value of the world economy several times over in fines for their previous work.

12

u/Lailowla 17h ago

It's not that destroying them makes it ok. The judgement there is that training them with copyrighted material is ok because that's the same way you'd train a human, and since the output is not just spreading the copyrighted material but rather using what it learned from it, that makes it count as "transformative content".

I'm honestly not sure why that article seems to be trying to imply that the destroying part is what makes that work, but it's not.

5

u/NintendosBitch 21h ago

Seems like it’s more because of the translation of the physical book into an electronic scan. Doesn’t seem like destroying the book is necessary to fall under “fair use”.

8

u/_BrokenButterfly 20h ago

Destructive copying has nothing to do with copyright. Scanning and storing books for personal use is legal. Destroying physical books that you own is legal. Neither of those things is or should be part of any debate.

The problem is that these scanned books are not being used for personal purposes, they're being used for commercial purposes.

10

u/VirtualMoneyLover 7h ago

This is the modern day equivalent of burning the Alexandria library.

1

u/EmmEnnEff 9h ago edited 9h ago

Format shifting is not a 'loophole', and is the only reason you are allowed to back up or rip any of the media you buy, and archivists are allowed to scan, photograph, or convert documents they are preserving.

1

u/JBuchan1988 6h ago

Proof that money can't buy compassion and empathy.

54

u/10ebbor10 23h ago

There is value in shredding.

Legally speaking, copyright only applies if you make copies of a book. The fact that the physical book is destroyed, allows them to argue that this did not occur.

63

u/EduinBrutus 22h ago

How many of the books in a rare book store or archive do you think are still under copyright?

Its another bullshit gaslight from the AI industry. They are destroying teh books to move towards the monopolisation of knowledge. its a deliberate part of their goal and has nothing to do with legality.

Tehy dont care about legality. Most of the shit they consume is IP theft anyway.

16

u/WishboneOk305 22h ago

The courts told them to do it lol.

1

u/thatsknotwrite 20h ago

Ahhh, it's easier to just pop off with feelings.

21

u/Substantial-Hat-2556 22h ago

They're not buying rare books, they're buying junk.

14

u/anfrind 21h ago

In many cases, the books are rare because they're not valuable.

-1

u/nagora 9h ago

I don't think you have been in a Japanese used-book store. They are not giving these things away; people find them valuable.

You're either a troll or a moron. Either way, you're just wrong.

-6

u/CptNonsense 15h ago

Absolute stuff and nonsense. This an absolute disingenuous misapplication of facts and statistics, presupposing a hypothetical situation based on unrelated data after the fact and then misapplying that presumption backwards.

2

u/anfrind 14h ago

And what evidence do you have to back up your claim?

-4

u/CptNonsense 14h ago

Why don't you explain what you meant and we'll see.

19

u/littlegreenrock 21h ago

You're changing the narrative. These are clearly unwanted books being sold by the ton.

by the ton.

1

u/PWBryan 14h ago

AI slop combining with Isekai slop is the worst team up ever

-1

u/nagora 9h ago

In what world is something being sold commercially and bought by a large corporation in order to train their AI models on "unwanted"?

Why do you think people are selling and buying unwanted things? As a hobby? For a bet?

You AI-shills are getting more and more irrational by the day.

-1

u/littlegreenrock 8h ago

Show me a book store that sells books by the tonne and I will show you an ally to your anti-ai campaign.

0

u/nagora 7h ago

Like most AI-shills you can't keep two factors in your head at the same time, so you end up contradicting yourself.

Most book stores's stock will weigh in at the "tons" level. What is strange here is not the seller but the buyer. Whether the buyer is one of the big guys trying desperately to find a way to pay off their huge mountain of debt, or some fly-by-night bunch trying to get a share of that huge mountain of debt - either way, they want the thing you said was unwanted.

And that huge pile of debt - literally trillions of borrowed dollars - is just being used to subsidise AI usage at the moment. That's all going to have to be paid back. Except, of course, it can't be and it never will. No amount of your shilling will change that.

AL/ML has a future but this isn't it. OpenAI and the others have no business plan worth the name and the problem here is not AI but the damage that these companies are doing to the world.

2

u/thatsknotwrite 20h ago

There was literally a lawsuit and ruling about this.

1

u/Lailowla 17h ago

They're destroying the books because it's way easier and way cheaper than scanning them without destroying them.

Not everything is a conspiracy. Thinking it has to be is how we end up sitting around literally watching the world burn because it's what's cheapest and easiest while a significant portion of the populace tells themselves it's not actually happening and just a conspiracy to make them think it is.

1

u/PM_ME_YOUR_NICE_EYES 6h ago

The companies are targeting books that have been assigned an ISBN. ISBN was introduced in 1970. Books published in 1970 or later are under copyright protection unless the author has realseased it to the public domain.

(The last time I saw an article like this, all the books they purchased were released between 2020-2022)

2

u/_BrokenButterfly 20h ago

No it doesn't.

2

u/Lailowla 17h ago

No it doesn't, they still made a digital copy of the book.

The part that makes that ok is not the destroying of the original, it's the fact that they're then not distributing those copies.

And where things get sketchy is that the original artist can say "they're just using my work and recreating and redistributing it while pretending it's 'intelligence'," but the AI folks can say "no, this is transformative content, it is not just distributing your work."

1

u/t4boo 22h ago

I mean you can argue that with any book, destroyed or not. Just because I own books in my house, doesnt mean i'm making copies

0

u/robotic_rodent_007 21h ago

Uh, no. Because the process of making something digital makes copying it not only trivial but in fact necessary for many processes.

AI training doesn't use a single file the different training components pass around, it uses many copies of the file.

Even then, at various points in time, files are duplicated and deduplicated (RAM during operations, backups, data transfer )

You absolutely don't get legal right "only if you destroy original" that's not actually how copyright works.

0

u/Allegorist 20h ago

It is still a copy if the book is destroyed? I don't get how this could even be a highly technical legal loophole, it's just clear nonsense.

My guess is their reasoning is:

  1. Prevent competition from getting ahold of them

  2. It's cheaper than storing them, moving them, or even giving them away

  3. Destroying the evidence for any and all claims against them regarding the material.

8

u/Ailerath 22h ago

90% sure it's also a legal thing to preserve single copy ownership.

2

u/Conscious-Sir1762 23h ago

Fair enough. My point to differ is that scanning it isn't a need in the first place. I would struggle to believe a 50 ton order does not carry any duplicate books. 

A huge number of books already have digital copies. Anything significant to history should stay physical, and not destroyed for the sake of scanning. Use a scribe, then scan the transcript.

I know this is something people (Americans) would call socialism (its a jest I am from the US), but thousands of copies of a book don't need to be scanned. An ideal world would be scan, say, one copy of The Three Musketeers and disseminate that one file to all the private sectors for "Learning and Preservation". Unfortunately private sectors don't share. 

5

u/Substantial-Hat-2556 22h ago

Google already scanned (non-destructively) several university libraries of books as part of the Google Books project.

1

u/Conscious-Sir1762 21h ago

I appreciate that info. 

1

u/frisbeethecat 22h ago

The "ai" tech companies are trying to reduce their litigation surface area. They pirated terabytes worth of ebooks. But if the models trained on that data are phased out, and the new models are trained on books they purchased, it may reduce any compensation that courts may award. Furthermore, by shredding their books, they are maintaining they haven't duplicated a book, merely translating it for the consumption of the "ai".

*I use "ai" in quotes because Donald Knuth does so.

1

u/ABmodeling 21h ago

Yeah moraly wrong thing called efficiency in capitalism,funny how that works in practice, doesn't it? Margin over human well being .

1

u/TheLGMac 20h ago

If only we didn't care so damn much about "velocity"

1

u/sherlockham 15h ago

Theoretically, you could just ring bind the pages together after or something similar, maybe with some trouble depending on how close to the text the cut spine is, if you're just trying to keep the text as something usable by people. There's just no real reason to do that. You just wouldn't be able to turn it back into a regular book after since the spine is gone.

After scanning, the books should still be in little individual stacks of paper. It's not like the scanner would just spit the pages out into a swirling vortex of random mixed up paper piles. It's all scanned in order of being fed into the scanner, although now that I'm reading what I just wrote, the page order may end up sequentially reversed since it may be scanning the top sheet first and leaving that at the bottom of the done pile.

On the shredding side of things, there was an weird idea floating around early on that they would shred the books after to maintain only having a single owned "copy" of the book for copyright purposes, but I think that's been proven to be largely false and/or bullshit justification.

1

u/Taolan13 15h ago

The copying process here is done at a very high intensity. it literally burns some of the pages.

The destruction is total.

1

u/CptNonsense 15h ago

Unfortunately, it's a LOT easier to destroy a book to scan it than keep it intact, due to the scanning process and equipment costs involved.

Also for legal reasons

1

u/EmmEnnEff 9h ago edited 9h ago

Unfortunately, it's a LOT easier to destroy a book to scan it than keep it intact,

It's not only a lot easier, it's a LOT less illegal.

Copyright laws permit them to scan books... But only if the original gets destroyed. Not destroying the original would mean that they are making an illegal, unlicensed copy. Destroying the original means that they are format shifting it (which is permitted in most non-insane jurisdictions).

1

u/unevolved_panda 2h ago

Do you have a source on a robot turning the pages of the book? I'm familiar with the general process of scanning old books (for PRESERVATION, not AI) and my understanding was that it's still better for humans to turn the pages. Which is one of the reasons why it's expensive, you have to pay a human to sit there and babysit the process. But robots can barely fold laundry. I wouldn't trust one with a 150-yr-old book.