Education7 min read

5 Ways Copyrighted Images Hide on Your Website (And How to Find Them)

Copyrighted images aren't always obvious. Learn 5 ways stock agencies and image-tracking services identify their images, and what you can check on your own site.

By Suraj, founder of PixGuard · Published · Updated

Cropping, resizing or filtering a stock photo doesn't stop its owner from recognizing it. Agencies and image-tracking services identify images by what they look like (visual fingerprints and image recognition), by the metadata inside the file, and by watermarks, some of which are easy to miss. Here are the 5 ways copyrighted images give themselves away, and what you can actually do about it.

1. Invisible Watermarks in Pixel Data (Steganography)

This one catches most people off guard because they've never even heard of it.

Some images carry invisible watermarks: patterns embedded directly into the pixel values. Every pixel in a digital image is just a set of numbers (0 to 255 for each color channel). If you change the last digit of those numbers by just 1, the image looks exactly the same to your eyes. But software that knows what to look for can read the pattern.

To put it simply: a pixel that's RGB(142, 87, 203) gets changed to RGB(143, 86, 202). You can't see the difference. Software absolutely can.

There are a few flavors of this:

  • LSB (Least Significant Bit) hides patterns in raw pixel values
  • DCT (Discrete Cosine Transform) embeds data in JPEG frequency coefficients, which is designed to hold up better under compression
  • DWT (Discrete Wavelet Transform) works at multiple resolution levels
  • Alpha channel data hides info in the transparency layer of PNG and WebP images

How well a hidden mark survives editing depends on the scheme. LSB marks usually break when an image is re-saved as a JPEG; frequency-domain marks are designed to be more robust. But invisible marks aren't the main way agencies find their images. Big-agency enforcement mainly relies on visual fingerprinting (method 3), which doesn't need any mark in the file at all.

2. EXIF Metadata That Sticks Around

Every photo taken with a digital camera or edited in professional software carries EXIF metadata. Think of it as a hidden label stitched into the file. This data can include:

  • A copyright notice (like "Copyright 2025 Getty Images")
  • The photographer's name
  • The camera serial number (yes, they can identify the exact camera)
  • GPS coordinates from where the photo was taken
  • What software was used to edit it (Photoshop, Lightroom, etc.)
  • ICC color profiles that indicate a professional editing workflow

Here's the thing: social media platforms usually strip EXIF data when you upload something. Regular websites vary. Depending on your CMS, image editor and CDN, the original file you uploaded may keep its metadata even if resized copies lose it. If you downloaded an image from a stock site and uploaded it straight to your site, that metadata may well have come along for the ride.

And even when EXIF gets stripped, the ICC color profile often hangs around. Seeing Adobe RGB 1998 or ProPhoto RGB in an image is a pretty strong signal that it came from a professional pipeline. It's a hint worth noting when you check an image by hand.

3. Perceptual Hashing (Fingerprinting)

Image-tracking systems create digital "fingerprints" of images using techniques such as perceptual hashing. Unlike a regular file checksum that changes if a single pixel is different, perceptual hashes produce similar results for images that look similar.

The algorithm works like this:

  1. Shrink the image down to a tiny thumbnail (like 32x32 pixels)
  2. Convert it to grayscale
  3. Calculate a hash based on the relative brightness of each area
  4. Compare that hash against a massive database of known copyrighted images

What this means in practice is that even if you resize the image, recompress it, adjust the colors, or crop it slightly, the perceptual hash often stays close enough to match. Image-tracking services such as PicScout, which Getty Images owns, rely on this kind of visual fingerprinting to find copies across the web.

Checking images one by one takes time

PixGuard flags images with copyright risk signals, so you know which ones to investigate first. Paste a page URL to check its 3 largest images, free and without signup.

4. AI Visual Recognition

This is the newest approach. Image-recognition models can detect similarity at a deeper level than just pixels.

What AI catches:

  • Images that have been heavily edited but still have the same composition
  • Screenshots that contain copyrighted photos somewhere in them
  • Photos used as backgrounds in collages or banners
  • Derivative works that are clearly based on a copyrighted original

These models understand what's in the image, not just what the pixels look like. So even if you've changed every single pixel through heavy editing, an image with clearly the same scene and composition can still be matched.

5. Logo and Text Watermarks You Didn't Notice

This sounds like it should be obvious, right? But it happens way more often than you'd expect. Stock agencies are sneaky about where they place their marks:

  • Semi transparent overlays across the whole image that disappear against certain backgrounds
  • Micro text hidden in busy areas like foliage, fabric textures, or crowd scenes
  • Corner logos that you cropped off, but left behind telltale artifacts
  • Single channel watermarks that only show up when you look at the red, green, or blue channel individually

A lot of people download preview images from stock sites, chop off the obvious watermark, and call it done. They can miss a second, fainter mark elsewhere in the image. And cropping out a watermark doesn't license the image; in the US, intentionally removing a watermark that identifies the owner, knowing it will help hide an infringement, can add separate liability under 17 U.S.C. 1202.

So How Do You Actually Detect All This?

Checking metadata, watermarks and fingerprint matches by hand, image by image, is slow. You'd need a whole toolkit of specialized software plus a database to compare against.

A scanner can take the first pass. With PixGuard, you can check one of your pages free: enter a page URL and we check the 3 largest images on that page for watermark patterns and other visual risk signals, up to 3 times a day. Uploads and 30 image scans need a free account. With an account, a website scan follows links across up to 50 pages and checks the images it finds, up to your plan's per-scan limit, for:

  • Copyright fields in EXIF, IPTC and XMP metadata
  • Visible watermark text and logos from 7 stock agencies, plus tiled watermark patterns
  • Matches against a database of reference image fingerprints
  • Signs of editing (error level analysis)
  • Signs of AI generation

Pro and Business add AI similarity matching (DINOv2 and SSCD), which can match edited or cropped copies, but only of images PixGuard has on file to compare against. PixGuard doesn't decode invisible watermarks. Each image gets a 0 to 100 risk score with the signals behind it, so you know which images to investigate first.

How to Stay Clean Going Forward

The best copyright problem is one you never have in the first place:

  1. Buy your stock photos properly. It's cheaper than you think.
  2. Save your license receipts. Match them to specific images.
  3. Use Creative Commons when it fits. Just follow the attribution requirements.
  4. Shoot your own stuff when you can. Nothing beats owning your content outright.
  5. Scan regularly. People add images to websites all the time, and not everyone checks the license first.

And remember: old images still count. In the US, a copyright claim generally has to be brought within 3 years (17 U.S.C. 507(b)), but when that clock starts depends on the circumstances, so an image someone uploaded three years ago can still lead to a demand letter today.

Want to know where your site stands? Scan your site with a free account: 30 image scans, valid for 30 days, no credit card required.

See which images on your site need a closer look

Paste a page URL and PixGuard checks its 3 largest images for copyright risk signals, free and without signup. A free account adds 30 image scans (valid 30 days); each site scan crawls up to 50 pages and checks up to 10 new images on the free plan.