DigiGuardiansDigiGuardians

Technology

Perceptual Hashing for Near-Duplicate Detection

A perceptual hash summarises what an image or frame looks like rather than its exact bytes, so re-encoded and resized copies still land close together. How the hashes are built and compared, and the edits that break them.

August 11, 20264 min read

A perceptual hash is a short signature computed from what an image looks like. Two images that look the same to a person produce signatures that are identical or nearly so, even when their files are completely different. That property makes perceptual hashing one of the cheapest ways to find near-duplicate copies of an image or a video frame.

Why a file hash fails after a re-encode

Cryptographic hashes are built for the opposite purpose. Change a single byte in a file and its cryptographic hash changes beyond recognition. That is exactly what you want when checking that a download has not been tampered with, and useless for finding pirated copies, because almost every copy has been re-encoded, resized, renamed or repackaged. A film compressed for a streaming site shares no cryptographic hash with the master, even though it looks the same on screen. The comparison of cryptographic and perceptual hashing sets the two side by side.

Cryptographic hashes still earn a place. Once a particular pirated file has been identified, its exact hash finds every byte-identical copy of that file across hosts. It just cannot connect the file back to the original work.

How a perceptual hash is built

The common methods share a pattern. The image is shrunk to a small greyscale version, discarding colour and fine detail, which are the things compression and resizing tend to alter. Something about the structure of that small image is then recorded as a sequence of bits.

The simplest methods compare each pixel with the average brightness, or with its neighbour, and write down a one or a zero. A widely used method applies a discrete cosine transform, from the same family of maths used in JPEG compression, and records whether each low-frequency component sits above or below the median. Low frequencies describe the broad layout of light and dark, which survives most everyday processing. Either way, the result is a compact string of bits that can be stored and compared cheaply in very large numbers.

Distance, thresholds and the cost of each mistake

Two perceptual hashes are compared by counting the positions where their bits differ, known as the Hamming distance. Images that look alike give a distance at or near zero. Unrelated images give a distance of roughly half the hash length, as if the bits were random.

The threshold sits somewhere in between. Set it tight and only near-identical copies match, which keeps false matches down but misses copies with heavier edits. Set it loose and more edited copies match, along with unrelated images that happen to share a layout: two dark posters with a bright title along the bottom, or two frames of empty sky. There is no correct threshold in the abstract. It depends on what a false match costs compared with a missed copy for that asset, and on whether a person will review the result anyway.

Edits that throw a hash off

Perceptual hashes tolerate the processing images go through in ordinary distribution: scaling, compression, mild colour and brightness changes, a little noise. They cope far less well with geometric edits. A horizontal mirror flips the layout and can move the hash a long way. A heavy crop changes what the shrunken image contains. Rotation, a large text overlay, a thick border or placing the image inside a bigger graphic all change the overall structure the hash summarises.

Some of this can be absorbed by hashing several variants of the reference, such as the mirrored version and common crops. Beyond that, edited copies need methods built on local features, which can match part of an image. Hashing is the fast first pass, and it works best when something sturdier stands behind it.

Hashing video frame by frame

For video, frames are sampled, often at shot changes, and each sample is hashed. A candidate video produces its own sequence of hashes, which is compared with the reference sequence. A run of close matches in the right order is strong evidence of a copy. A scattering of isolated matches may be a trailer, a review or coincidence. This is one of the simplest forms of video fingerprinting, and it handles the straightforward reuploads that make up much of what monitoring turns up.

Where hashing sits in the queue

Hashing is fast and cheap enough to run on everything collected. Its role is to rank candidates and tie each one to a specific reference asset, so that reviewers start with the likeliest copies and know what each was matched against. A close hash is a strong lead. The decision to send a notice still depends on the page: whether the use is authorised, whether it is the rights holder's own marketing, and what exactly is on offer.

  • Hash Matching
  • Technology

Keep reading.

Piracy moves fast. Takedown should move faster.

Tell us what you protect. We'll map where your titles leak and show you what we'd remove first.

First report free · 14-day trial · No obligation

Stay ahead of the pirates.

No spam, just the takedowns, threats and reports worth your inbox.