DigiGuardiansDigiGuardians

Technology

Web Scraping in Copyright Monitoring

Scraping is the extraction half of copyright monitoring: pulling titles, players, hosts and download links out of pages already found. How it follows a listing to the file, captures evidence and re-checks after a notice.

August 11, 20264 min read

Web scraping and web crawling are often used as if they meant the same thing. In copyright monitoring they do different work. A crawler finds pages. A scraper reads a page that has already been found and extracts specific facts from it: which title is offered, which episode, which player, which file host, which download links. Scraping turns a URL in a queue into a structured record that an analyst, or an enforcement process, can act on.

Pirate sites are templates

Most pirate sites are built from a few page types. A streaming index has a listing page per title, an episode page per episode, and a player area that loads a stream from somewhere else. A link aggregator has a table of hosts for each release. Once the template of a site is understood, a scraper can read every page built on it and pull out the same fields each time.

Templates change, by accident or on purpose. Operators rename page elements, move content into scripts, or rotate layouts specifically to break automated extraction. A scraping setup has to notice when a site that used to parse cleanly starts returning empty fields, and flag that site for repair instead of quietly reporting that nothing was found. Silent failure is the most expensive kind, because it looks exactly like success.

Following a listing to the file

The page a user lands on is rarely where the infringing content lives. Take a hypothetical chain. A listing page shows the poster and an episode list. The episode page embeds a player in an iframe served from a second domain. That player requests a playlist from a third domain, and the playlist points to video segments on a content delivery network. Download links on the same page go through a URL shortener and an advertising interstitial before they reach a cyberlocker.

Every hop matters for enforcement. The listing site, the player host, the segment host and the cyberlocker are often different operators with different hosting providers. Removing a link on the listing page leaves the file where it is. Removing the file at the cyberlocker breaks every listing that points to it. Scraping that follows the chain to the end is what makes it possible to act on the source as well as the link, and the comparison of source takedown and link removal sets out why both matter.

Following the chain often means rendering the page as a browser would, because players are assembled by scripts after the page loads. Some sites only reveal the player after a click on a play button or a server selector, and the scraper has to reproduce that step faithfully to see what a viewer would see.

Evidence at the moment of detection

Pirate pages change fast. A link that was live when found may be dead an hour later and replaced by a new one. A notice filed against a page that has since changed can be rejected, and a dispute months later will turn on what the page showed at the time.

Extraction should therefore capture evidence alongside data:

  • a screenshot of the page as rendered
  • the relevant page source
  • the resolved URL of each hop in the chain
  • the status and headers the servers returned
  • a precise timestamp for each of the above

Stored together, that record shows what was offered, where and when, without anyone having to revisit the site.

Checking again after the notice

Scraping has a second job once enforcement starts. Re-checking a URL tells the rights holder whether a notice worked, and the possible outcomes are more varied than up or down. The file may be gone while the listing stays live. The listing may now point to a new host. The page may return an error in some regions only. The content may have been swapped for a different title. Each outcome suggests a different follow-up, and recording them over time shows which hosts and platforms comply and which need escalation.

The line monitoring should not cross

Copyright monitoring collects what any member of the public can see. It should not log into accounts it has no right to use, should not get around access controls, and should not request pages at a rate that degrades a site. Personal data that appears on pages, such as usernames in comment threads, should be collected only where the case needs it and handled under the applicable privacy rules, which differ between jurisdictions. A monitoring programme should have its collection rules written down rather than left to the judgement of each script.

  • Crawlers
  • Technology

Keep reading.

Piracy moves fast. Takedown should move faster.

Tell us what you protect. We'll map where your titles leak and show you what we'd remove first.

First report free · 14-day trial · No obligation

Stay ahead of the pirates.

No spam, just the takedowns, threats and reports worth your inbox.