Knowledge
How Infringing Link Discovery Works
A pirate link is usually a chain of pages, redirects and hosts. How link discovery gathers candidates, resolves each chain to its source and decides which links are worth a notice.
An infringing link is rarely a single address. A visitor clicks a title on an index page, lands on a page with a player or a download button, passes through a link shortener and one or more advertising interstitials, and finally reaches a file on a cyberlocker or a stream served from a video host. Link discovery is the work of finding those chains, resolving them to their end points, and deciding which links in each chain can be acted on.
Where candidate links come from
Discovery runs on several feeds at once, because pirate audiences arrive by different routes:
- Search engines, using title and access-word queries in each market.
- Link aggregators and index sites that catalogue titles and point to external hosts.
- Forums and community boards where releases are posted with file host links.
- Social platforms, where links sit in posts, comments, bios and replies.
- Messaging channels and groups that post links, files or invitations to other channels.
- Embedded players inside otherwise ordinary pages, such as a blog post with a stream.
Each feed produces candidate URLs. Many are duplicates, dead, irrelevant or legitimate, which is expected. The goal at this stage is coverage. Filtering comes next.
Resolving the chain
A candidate URL is resolved by following it the way a user would. The analyst or tool loads the page, identifies what it offers, follows outbound links and buttons, records each redirect and shortener, and continues until reaching content or a dead end. The result is a chain with typed layers: the index page, the landing page, any intermediate redirects, and the hosting end point.
Resolution matters because the layers belong to different services with different obligations. An index page that only links may be handled by the site's host or by search engines. The file on a cyberlocker can be removed by the locker. A shortener provider can disable a link that points to infringing content. Without resolution, a programme reports whatever it found first, which is usually the least durable layer.
Resolution also exposes traps. Some chains end in malware downloads, fake players that demand a sign-up, or survey pages, with no copy of the work at all. Those pages still exploit the title, but they need a different response from a working copy.
Classifying and deduplicating
Each resolved link is classified: which work it relates to, which version or episode, what kind of service hosts it, whether it is working at the time of capture, and whether it falls under the whitelist. Duplicates are merged, since the same locker file may be linked from many pages. This is where a programme's view becomes useful. Instead of a long flat list of URLs, the team sees that a small number of source files are feeding most of the visible links for a title.
That picture drives priorities. Removing one source file that many pages point to does more than removing those pages one by one, although both may be needed.
What makes a link reportable
A link can be reported when it is specific, identifies the work, and is shown to be infringing at the time of capture. In practice the record holds the exact URL, the title and version it offers, a capture of the page and, where possible, of the content playing or downloading, the time of capture and the observed host. A search result pointing to a page with a dead player, or a page that names a film it does not carry, needs different handling from a page with a working copy.
Every link is verified by a person before a notice is filed. Automated discovery is good at volume and poor at context: a review embedding the official trailer, a licensed partner's page and a news story about a leak can all look like matches.
Choosing which links to act on
Link discovery feeds two kinds of action. Source actions remove the file or stream from the host. Link actions remove the page, post or search result that points to it. The comparison of source takedown and link removal sets out when each is enough. Most programmes need both: source removal disables every link to that file, while delisting reduces how many people find the pages and catches pages that would otherwise swap in a new file.
The structure of pirate distribution that these chains belong to is covered in how online piracy works. DigiGuardians acts on the source as well as the link, and every action is documented so the client can see what was found, where, and what happened to it.
- Detection
- Knowledge


