Lawsuit Accuses Anna's Archive of Hacking WorldCat, Stealing 2.2 TB Data

ancuuiqter@lemmy.world · edit-2 1 year ago

Lawsuit Accuses Anna's Archive of Hacking WorldCat, Stealing 2.2 TB Data

EinatYahav@lemmy.today · 1 year ago

tear down every paywall

body_by_make@lemmy.dbzer0.com · edit-2 1 year ago

Yes, let only the rich control your thoughts.

I’m not surprised this will get downvoted here, I’m as much of a pirate as anyone, but news needs to be paid or only people who can afford to control the news without income will control the news.

ShepherdPie@midwest.social · edit-2 1 year ago

Npt saying you’re right or wrong but paid news has been the model for quite a while now and that has resulted in 24 hour talking heads on TV, paid stories, clickbait, and people resorting to word of mouth on places like Facebook for all their news. It’s not as if the current trajectory is any better than your hypothetical one.

dangblingus@lemmy.dbzer0.com · 1 year ago

Having a copy of the data publicly available through Anna’s Archive is a direct threat to its business.

How would it hurt WorldCat’s business given that the service they offer is free? If the information, that being the location of books and articles in specific libraries around the USA, was freely available on another site, what value has been lost?

RogueBanana@lemmy.zip · edit-2 1 year ago

Just going by anna’s blog post their business model seems to be trading information ie sharing the full database of hundreds of millions records with their memeber’s own records so the list keeps growing as more members join. Although I don’t see why they need a monopoly on said information given any other library would still continue working with them for their free streamlined process. There could be more to it but feels like they are wasting resources on this instead of putting them in things that actually matter.

Edit: also I don’t think they scrapped or have information about the members like location of each book, simply just the metadata so it really seems harmless to me

xiao@sh.itjust.works · 1 year ago

Wish AA gonna be fine, they made me save literally hundred of US dollars…

ancuuiqter@lemmy.world · edit-2 1 year ago

The official Anna’s Archive Reddit account, AnnaArchivist, has responded to an r/Annas_Archive post linking the same Torrent Freak article:

Thanks! We’re not making any public statements about this lawsuit but rest assured we’re fine.

EveryMuffinIsNowEncrypted@lemmy.blahaj.zone · 1 year ago

Sigh… Whelp, time to go download a shit-ton of stuff before yet another friendly port goes down…

Snot Flickerman@lemmy.blahaj.zone · 1 year ago

https://annas-blog.org/worldcat-scrape.html

Relevant blog post. AA knew the risks in this, and this is sort of expected.

Darkassassin07@lemmy.ca · edit-2 1 year ago

Gotta wonder what their plan is. The lawsuit was an obvious outcome, and they haven’t exactly made much effort to make their actions appear legal.

I don’t see AA winning this one. Data’s out there though; no taking that back. Maybe they’ve just accepted the consequences… A martyr as it were.

BarrierWithAshes@kbin.social · 1 year ago

AA’s based outta Kazakhstan though. Lotta good a lawsuit filed in Ohio’s gonna do. At most I could see American ISPs implementing a DNS-level block against the site.

Darkassassin07@lemmy.ca · 1 year ago

Oh. Lol, get fucked WorldCat.

ancuuiqter@lemmy.world · 1 year ago

Would you be able to share where you learned that Anna’s Archive is based in Kazakhstan?

BarrierWithAshes@kbin.social · 1 year ago

I remember reading it on the site but I cannot find it now. I know for a fact she is based in Kazakhstan. So says her wikipedia page.

ancuuiqter@lemmy.world · edit-2 1 year ago

Maybe you’re thinking of Sci-Hub and its founder, Alexandra Asanovna Elbakyan?

I could not find a location on Anna’s Archive’s wiki page.

MotoAsh@lemmy.world · edit-2 1 year ago

I mean… it’ll all come down to how they accessed the data. If they had a public portal and no EULA, they can push rocks. If the data wasn’t public or the ‘theives’ had to use non-standard channels, or otherwise violated an EULA, they’re likely screwed. Especially if they had to go through abnormal channels.

I know their data can be accessed publicly, but I’m pretty sure it’s under license. You cannot just use any old thing found in public… That’s the biggest reasons the AI models are technically theft: they weren’t licensed to commercially profit off of 99.99% of the things their LLMs are trained on, but the law and politicians are WAY behind the times. Commercial data they’d normally have to pay for is suddenly magically OK when laundered through an LLM…

Dkarma@lemmy.world · 1 year ago

“AI models are technically theft: they weren’t licensed to commercially profit off of 99.99%”

This is simply a lie. There is no license like what you describe. You never need a license to view or learn from something given away completely free on the internet. You guys keep pretending there’s a law that says otherwise . There is not or you’d post it.

Copyright does not cover viewing or experiencing a piece.

MotoAsh@lemmy.world · edit-2 1 year ago

Notice how I said “commercially profit” too. Read all the words next time.

Also LLMs do not “learn” anything, you idiot. That’s the entire point. They mathematically blender things. They DO NOT learn and create.

ancuuiqter@lemmy.world · 1 year ago

Here are the court filings if anyone would like to read them:

https://archive.org/details/gov.uscourts.ohsd.287709/

The following is a link to the docket (which the above link draws from), so people can follow the progress of the lawsuit:

https://www.courtlistener.com/docket/68157923/oclc-online-computer-library-center-inc-v-annas-archive/

ancuuiqter@lemmy.world · 1 year ago

As to how Anna’s Archive accomplished their data scraping, this is what OCLC is claiming (see page 62-63):

These attacks were accomplished with bots (automated software applications) that “scraped” and harvested data from WorldCat.org and other WorldCat®-based research sites and that called or pinged the server directly. These bots were initially masked to appear as legitimate search engine bots from Bing or Google.

To scrape or harvest the data on WorldCat.org, the bots searched WorldCat.org results, running a script based on OCN for individual JavaScript Object Notation, or “JSON,” records. As a result, WorldCat® data including freely accessible and enriched data, such as OCNs, were scraped from individual results on WorldCat.org.

The bots also harvested data from WorldCat.org by pretending to be an internet browser, directly calling or “pinging” OCLC’s servers, and bypassing the search, or user interface, of WorldCat.org. More robust WorldCat® data was harvested directly from OCLC’s servers, including enriched data not available through the WorldCat.org user interface.

Finally, WorldCat® data was harvested from a member’s website incorporating WorldCat® Discovery Services, a subscription-based variation of WorldCat.org that is available only to a member’s patrons. Again, the hacker pinged OCLC’s servers to harvest WorldCat® records directly from the servers. To do this through WorldCat® Discovery Services/FirstSearch, the hacker obtained and used the member’s credentials to authenticate the requests to the server as a member library.

From WorldCat® Discovery Services, hackers harvested 2 million richer WorldCat® records that included data not available in WorldCat.org. This hacking method resulted in the harvesting of some of OCLC’s most proprietary fields of WorldCat® data.

These hacking attacks materially affected OCLC’s production systems and servers, requiring around-the-clock efforts from November 2022 to March 2023 to attempt to limit service outages and maintain the production systems’ performance for customers. To respond to these ongoing attacks, OCLC spent over 1.4 million dollars on its systems’ infrastructure and devoted nearly 10,000 employee hours to the same.

Despite OCLC’s best efforts, OCLC’s customers experienced many significant disruptions in paid services during the aforementioned period as a result of the attacks on WorldCat.org, requiring OCLC to create system workarounds to ensure services functioned.

During this time, customers threatened and likely did cancel their products and services with OCLC due to these disruptions.

Because OCLC had to combat these persistent hacking attacks, OCLC was forced to divert existing personnel and resources from OCLC’s other products and services. As a result, OCLC’s development and improvements to other products and services were delayed and limited.

OCLC has devoted, at various times, ten or more employees to respond to and mitigate the harm from these attacks from October 2022 to present.

conciselyverbose@kbin.social · 1 year ago

None of this is “hacking”

Lawsuit Accuses Anna's Archive of Hacking WorldCat, Stealing 2.2 TB Data

Lawsuit Accuses Anna's Archive of Hacking WorldCat, Stealing 2.2 TB Data

Lawsuit Accuses Anna's Archive of Hacking WorldCat, Stealing 2.2 TB Data * TorrentFreak