How Getty Images finds websites using unlicensed photos
Getty Images uses PicScout perceptual-hash crawling to detect unlicensed images across the web. Learn how the system works and how to reduce your exposure.
By Suraj, founder of PixGuard · Published
Getty Images uses an automated system called PicScout to crawl billions of web pages and identify images that match its licensed catalog, even when those images are cropped, resized, or color-adjusted. The process is largely automated and operates at a scale no human reviewer could achieve. Understanding how the detection works helps you assess your own risk and take steps before a demand letter arrives.
What PicScout is and how it works
PicScout is a computer vision company that Getty Images acquired in 2011. Its core technology is perceptual hashing, sometimes called image fingerprinting. A perceptual hash is a compact numerical representation of an image's visual structure: the distribution of edges, tones, and shapes across the frame. Unlike a cryptographic hash (which changes completely if you alter even one pixel), a perceptual hash remains similar even after cropping, scaling, compression, or moderate color adjustment.
PicScout generates perceptual hashes for every image in Getty's catalog. A crawler then visits web pages, downloads images it finds, computes their perceptual hashes, and compares those hashes against Getty's database. A close enough match triggers a flag for human review and, ultimately, a demand letter.
Because perceptual hashing is resilient to common image edits, techniques that once helped people avoid detection (resizing an image, adding a faint border, adjusting brightness) are largely ineffective against the current generation of matching systems. The fingerprint survives most non-destructive edits.
How extensive is the crawl?
PicScout's crawling operation is continuous, not a periodic one-time audit. It uses infrastructure similar to a search engine spider, following links across publicly accessible web pages. Any page that is publicly accessible can be crawled. This includes:
- Blog posts and news articles
- E-commerce product listings and category pages
- Restaurant and small business websites
- Real estate listing pages
- Social media profiles and posts on platforms that allow public access
- Press releases and online newsrooms
- Older or archived pages that remain indexed by search engines
Pages behind a login, a paywall, or a robots.txt disallow directive may be harder to crawl, but they are not guaranteed to be invisible. The robots.txt standard is advisory rather than technically enforced.
What happens after a match is detected
When PicScout's system flags a potential match, a reviewer verifies whether the image was licensed. Getty maintains a customer database and can cross-reference whether the domain in question has ever purchased a valid license for the image.
If no valid license exists, the case is forwarded to Getty's in-house enforcement team or to an external enforcement partner. Getty has worked with firms including PicRights, and has also brought direct infringement suits. The initial contact is almost always a demand letter requesting retroactive licensing fees, often calculated at a multiple of what a license would have cost at the time of use.
How other agencies use similar technology
Getty is the largest and most well-known agency using automated detection, but it is not the only one. Shutterstock, Adobe Stock, and smaller regional agencies use comparable systems. Some agencies outsource enforcement entirely to firms like PicRights, which pays the agency for the right to pursue infringement on its behalf and keeps a portion of any collected settlement.
This means you can receive a demand letter from a firm you have never heard of, over an image from an agency you did not know you were using. The enforcement firm acting as claimant is a legally valid arrangement: they hold an assignment or exclusive license of the enforcement rights.
Checking images one by one takes time
PixGuard flags images with copyright risk signals, so you know which ones to investigate first. Paste a page URL to check its 3 largest images, free and without signup.
What makes an image higher risk for detection
Not all unlicensed images carry equal detection risk. Factors that increase the likelihood of detection include:
High-traffic pages. Well-linked pages receive more crawl priority. A Getty image on your homepage is more likely to be found than one buried in a 2017 blog post, though that older post is not invisible.
Image resolution close to the original. Highly compressed or very small thumbnails produce weaker perceptual hash matches, though this is not reliable protection.
Images from large agency catalogs. Getty's catalog includes hundreds of millions of images. Images from this catalog are well-indexed for detection.
Long display history. The longer an image remains on a public page, the more crawl cycles it is exposed to. A blog post from 2015 that you forgot about may have been accumulating exposure for years before the detection system flags it.
This is why auditing older content matters as much as reviewing new posts.
The role of metadata and watermarks in enforcement
Beyond perceptual hashing, Getty images typically contain IPTC metadata fields that identify the image's source, rights holder, and copyright notice. If these fields are present in an image file on your server, they serve as additional evidence of the image's origin even in cases where the visual match alone might be borderline.
Some Getty images also carry invisible digital watermarks encoded by systems like Digimarc, which embed a signal into the image data that survives many edits and can be decoded without access to the original file. These watermarks are not visible to the human eye but are detectable by software.
PixGuard checks images for both metadata signals and detectable watermark indicators. Running a scan at PixGuard before you publish helps surface images that may carry these identifiers. The watermark detection tool specifically flags invisible watermarks that would be missed in a visual review.
Why "I did not know" rarely ends the matter
A common response when receiving a Getty demand letter is to explain that the image was used accidentally or without knowledge of its protected status. US copyright law does recognize innocent infringement as a mitigating factor: statutory damages can be reduced at a court's discretion when the infringer had no reason to believe the use was infringing. However, this reduction is discretionary, not automatic, and applies primarily through court proceedings rather than demand letter negotiations. In practice, enforcement firms rarely offer innocent infringement rates in pre-litigation settlements.
For detailed guidance on how to respond once a Getty letter arrives, see our Getty Images demand letter guide.
Steps to reduce your exposure now
- Audit your site for images you did not photograph yourself or license through a documented transaction. Do not overlook older blog posts and archive pages.
- Verify that any stock images you are currently using have a valid, in-scope license for your use case.
- Remove or replace any images you cannot confirm are properly licensed.
- For future content, use only images where you can document the license: your own photos, confirmed-license stock purchases, or free-platform downloads with a saved record of the download date and source URL.
- Run your site through PixGuard to identify images carrying stock agency fingerprints or watermark signals before the crawlers find them first.
Comparison: what detection methods can and cannot catch
| Detection method | Catches cropped images | Catches color-adjusted images | Catches resized images | Defeats with screenshot |
|---|---|---|---|---|
| Pixel-exact hash match | No | No | No | Partial |
| Perceptual hash (PicScout) | Yes | Yes | Yes | Partial |
| IPTC / EXIF metadata | Depends on tool | Yes if metadata intact | Yes if metadata intact | Usually no |
| Invisible watermark (Digimarc) | Yes | Yes | Yes | Usually yes |
| Visible watermark (removed) | Not applicable | Not applicable | Not applicable | Not applicable |
Frequently Asked Questions
Can Getty detect images that I cropped or resized? Yes. PicScout's perceptual hashing is specifically designed to match images that have been cropped, resized, compressed, or color-adjusted. Minor edits do not meaningfully reduce the match probability.
Does Getty crawl social media accounts? Getty does monitor some social media platforms, particularly public business profiles and pages. The level of coverage varies by platform due to access restrictions and each platform's robots policy.
If I am using an image under a Creative Commons license, can Getty still send me a demand letter? It can send a letter, but if you have documentation that the image is legitimately licensed under Creative Commons and that Getty is not the rights holder, that is a strong defense. Verify that the Creative Commons license is actually associated with that specific image and that it was in effect when you downloaded it.
Will removing the image stop a demand? Removal stops ongoing infringement but does not eliminate past infringement. Getty can still pursue damages for the period during which the unlicensed image was publicly displayed. Removal is still the right first step.
How long does Getty typically wait before sending a demand? There is no fixed timeline. Detection can trigger a demand months or years after the image was first published. US copyright infringement claims have a three-year statute of limitations measured from when the rights holder knew or should have known about the infringement.
The most effective defense against automated detection is making sure the images on your site are ones you own or licensed properly. Run a free scan at PixGuard to identify images on your site that carry stock agency fingerprints or watermark signals before the crawlers find them first.
See which images on your site need a closer look
Paste a page URL and PixGuard checks its 3 largest images for copyright risk signals, free and without signup. A free account adds 30 image scans (valid 30 days); each site scan crawls up to 50 pages and checks up to 10 new images on the free plan.