# Online shadow libraries – and why they’re being…

> Listen to expert on shadow libraries Balázs Bodó on The Conversation Weekly podcast.

*Section: Technology — By The Conversation — Published August 7, 2026 — 2 min read*

Canonical URL: https://dailyjunction.org/technology/tc-online-shadow-libraries-and-why-they-re-being-implicated-in-ai-copyright-lawsuits
Tags: the conversation

<div class="theconversation-article-body">
    <figure>
      <img alt="A computer screen with an image of a library on it. " src="https://images.theconversation.com/files/750650/original/file-20260728-102-b1092u.jpg?ixlib=rb-4.1.1&rect=0%2C0%2C3840%2C2160&q=45&auto=format&w=754&fit=clip" />
        <figcaption>
          
          <span class="attribution"><a class="source" href="https://www.shutterstock.com/image-photo/convenient-library-computer-770095483?trackingId=f08ed0e6-752b-42a4-ab34-c45cc5240f15&listId=searchResults">won gou choi/Shutterstock</a></span>
        </figcaption>
    </figure>

  <span><a href="https://theconversation.com/uk/team#gemma-ware">Gemma Ware</a>, <em><a href="https://theconversation.com/">The Conversation</a></em></span>

  <div style="width: 100%; height: 200px; margin-bottom: 20px; border-radius: 6px; overflow: hidden;">

<iframe style="width: 100%; height: 200px;" frameborder="no" scrolling="no" allow="clipboard-write" seamless="" src="https://player.captivate.fm/episode/4291d2d4-1a24-4294-8bfe-2978642a8153/" width="100%" height="400"></iframe>

</div>

<p></p>

<p><iframe id="tc-infographic-561" class="tc-infographic" height="100" src="https://cdn.theconversation.com/infographics/561/4fbbd099d631750693d02bac632430b71b37cd5f/site/index.html" width="100%" style="border: none" frameborder="0"></iframe></p>

<p>A few days before Christmas 2025, a secretive online activist group called Anna’s Archive said it had downloaded roughly 86 million audio files and  256 million rows of track metadata from Spotify’s catalogue. </p>

<p> Anna’s Archive is a shadow library. It  holds one of the world’s largest online collection of pirated books, academic papers and now music. Soon, batches of the files, <a href="https://cyberinsider.com/annas-archive-releases-massive-300tb-spotify-music-scrape/">300 terrabytes in total</a>, started circulating on torrent file-sharing websites. </p>

<p>By April, a district judge in New York had ordered Anna’s Archive to <a href="https://www.musicbusinessworldwide.com/spotify-and-record-labels-win-322m-default-judgment-against-pirate-site-annas-archive/">pay US$322 million dollars in damages to Spotify</a> and three major record labels in a copyright infringement case over the hack. Then following month, the same judge ordered <a href="https://www.publishersweekly.com/pw/by-topic/digital/copyright/article/100463-court-rules-against-anna-s-archive-in-copyright-lawsuit.html">Anna’s Archive pay $19.5 million</a> in damages to a group of 13 major publishers for illegally copying and distributing works they’ve published. The judgement also ordered internet providers to block access to the site.</p>

<p>But the people behind Anna’s Archive didn’t show up in court, and these rulings will be very difficult to enforce. </p>

<p>In this episode of <a href="https://theconversation.com/topics/the-conversation-weekly-98901">The Conversation Weekly</a> podcast, Balázs Bodó, a professor of information law and policy at the University of Amsterdam in the Netherlands, takes us inside the history of these shadow libraries to explain how thousands of copies of books and academic articles ended up in these online archives. </p>

<blockquote>
<p>Anna’s Archive, is just the latest iteration of this endless stream of people who from since the beginning of time, try to archive everything, says Bodó.</p>

<p>During the last 20 to 30 years, there have been countless efforts to build digital libraries. Shadow libraries, we call them, because they are text collections, but they are not sanctioned officially.</p>
</blockquote>

<p>With a number of AI companies currently facing copyright lawsuits over allegations they’ve used shadow libraries to train their large language models, we ask what this means for the future of these online archives.  </p>

<div style="width: 100%; height: 200px; margin-bottom: 20px; border-radius: 6px; overflow: hidden;">

<iframe style="width: 100%; height: 200px;" frameborder="no" scrolling="no" allow="clipboard-write" seamless="" src="https://player.captivate.fm/episode/4291d2d4-1a24-4294-8bfe-2978642a8153/" width="100%" height="400"></iframe>

</div>

<p><em>Listen to Bodo on <a href="https://pod.link/1550643487">The Conversation Weekly</a> podcast.</em> </p>

<h2>Disclosure statement</h2>

<p><em>Balázs Bodó has received ERC funding. He was the project lead for Creative Commons Hungary and a member of Hungary’s National Copyright Expert Group. He has advised several public and private institutions on digital archives, content distribution, online communities, business development.</em> </p>

<h2>Credits</h2>

<p><em>This episode of The Conversation Weekly was written and produced by Gemma Ware and Mend Mariwany with editing help from Ashlynne McGhee. Mixing by Michelle Macklem and theme music by Neeta Sarl.</em></p>

<p><em>Newsclips in this episode from <a href="https://www.youtube.com/watch?v=wRtiMtl9bCw">India Today</a>, <a href="https://www.youtube.com/watch?v=Gb9TJMDNiM4">CBS News</a>, <a href="https://www.youtube.com/watch?v=o4Pgirl3zAM">CBS Miami</a> and <a href="https://www.youtube.com/watch?v=tVpIt0Q35uc">NBC News</a>.</em></p>

<p><em>Listen to The Conversation Weekly via any of the apps listed above, download it directly via our <a href="https://feeds.captivate.fm/the-conversation-weekly/">RSS feed</a> or find out <a href="https://theconversation.com/how-to-listen-to-the-conversations-podcasts-154131">how else to listen here</a>. A transcript of this episode is available via the Apple Podcasts or Spotify apps.</em><!-- Below is The Conversation's page counter tag. Please DO NOT REMOVE. --><img src="https://counter.theconversation.com/content/288557/count.gif?distributor=republish-lightbox-basic" alt="The Conversation" width="1" height="1" style="border: none !important; box-shadow: none !important; margin: 0 !important; max-height: 1px !important; max-width: 1px !important; min-height: 1px !important; min-width: 1px !important; opacity: 0 !important; outline: none !important; padding: 0 !important" referrerpolicy="no-referrer-when-downgrade" /><!-- End of code. If you don't see any code above, please get new code from the Advanced tab after you click the republish button. The page counter does not collect any personal data. More info: https://theconversation.com/republishing-guidelines --></p>

  <p><span><a href="https://theconversation.com/uk/team#gemma-ware">Gemma Ware</a>, Head of Audio, The Conversation UK, <em><a href="https://theconversation.com/">The Conversation</a></em></span></p>

  <p>This article is republished from <a href="https://theconversation.com">The Conversation</a> under a Creative Commons license. Read the <a href="https://theconversation.com/online-shadow-libraries-and-why-theyre-being-implicated-in-ai-copyright-lawsuits-288557">original article</a>.</p>
</div>

---
Daily Junction — https://dailyjunction.org/technology/tc-online-shadow-libraries-and-why-they-re-being-implicated-in-ai-copyright-lawsuits
