Skip to main content
← back to blog
Privacy

How to stop AI from scanning and tracking your photos (what works, what doesn't)

Training scrapers, face recognition, and platform AI all feed on shared photos. Here's every real opt-out and protection available in 2026 — and an honest account of their limits.

  • AI
  • privacy
  • do-not-train
  • opt-out
  • tracking

People are asking two versions of this question. One: how do I stop AI companies from training on my photos? Two: how do I stop my photos from being used to track and identify me? They overlap less than you'd think, the tools for each are different, and most advice online quietly ignores how many of those tools are voluntary.

Here's the honest map of what you can actually do in 2026.

First, what's actually happening to shared photos

Three separate things, worth keeping apart:

  • Training scrapes. Crawlers collect public web images into datasets that train image generators and vision models. If a photo is publicly reachable, assume it's been collected at least once.
  • Face recognition. Services have built searchable face indexes from scraped public photos. Upload-a-face, find-the-person. Several have lost lawsuits over it; the indexes exist regardless.
  • Metadata tracking. Separate from AI entirely: the EXIF in a shared file — GPS, timestamps, device identifiers — reveals where you were and when, no machine learning required.

Each has different defenses. No single tool covers all three — anyone telling you otherwise is selling something.

The opt-out signals (voluntary, but improving)

A family of machine-readable "don't use this" signals now exists for images:

IPTC's Data Mining field. The photo-metadata standard added a dedicated property for prohibiting data mining and AI/ML training. You embed it in the file itself; compliant crawlers read it and skip the image. Some major crawlers say they respect it. You can add it to any JPEG, PNG, or WebP right now with our do-not-train tagger — it writes the IPTC property and your copyright in the browser, without touching a pixel or uploading the file.

C2PA's do-not-train assertion. Content Credentials can carry a signed "do not use for training" assertion — Adobe tools can write it, and it's part of the same provenance stack the EU now requires AI providers to participate in.

Site-level signals. If you control the website, robots.txt rules for AI crawlers (most major AI companies now publish their crawler names) and noai meta tags cover everything you host.

The honest caveat, and it's a big one: all of these are requests, not locks. Compliant crawlers honor them; scrapers that don't care, don't. They're worth using — they cover the mainstream collectors, and they establish your objection in a way that increasingly matters legally — but they are etiquette with growing teeth, not enforcement.

There's also a structural catch we've covered before: most platforms strip metadata on upload. An IPTC do-not-train flag inside your photo doesn't survive posting to Instagram — the platform discards it along with everything else. In-file opt-outs work best for images you host yourself or share as files.

Platform opt-outs (use them, they're real)

Where your photos live on a platform, the platform's own AI settings matter more than anything in the file:

  • Meta (Instagram/Facebook): has an objection process for using your content to train Meta AI — stronger in the EU/UK, where objections are honored under GDPR pressure.
  • X: a settings toggle controls whether your posts (including images) feed Grok training.
  • LinkedIn, TikTok, and others: training toggles have appeared in privacy settings across 2024–2026; check yours after every major terms update, because defaults favor the platform.

These are contractual rather than technical — but unlike scraper etiquette, platforms can be held to their own settings.

Adversarial protection (for artists, mostly)

Tools like Glaze and Nightshade alter your image's pixels so AI models mis-learn from it — style cloaking and data poisoning respectively. They're the only technical (rather than voluntary) anti-training measure that exists, and they're aimed at artists protecting a body of work. Know the tradeoffs: subtle visual artifacts, an arms race with model builders, and no protection for photos already scraped.

What metadata removal does and doesn't do here

Let's place our own tool honestly on this map.

Stripping metadata does not stop AI scanning. Scrapers collect pixels; a clean file scrapes exactly as well as a dirty one. Anyone who implies EXIF removal protects you from AI training is wrong.

What it does is limit what any scan learns about you. A scraped photo with intact metadata hands over your home coordinates, your daily patterns, your device identity, and your editing stack alongside your face. The same photo cleaned hands over pixels. If your photos are going to end up in datasets and indexes you can't control — and public ones will — the difference between "my face" and "my face, home address, and daily schedule" is the part you still control.

That's the realistic posture: assume scanning, minimize what scanning yields. Check what a photo is carrying with CleanImages before it goes anywhere — GPS, device info, timestamps, AI tags — and strip what you don't want traveling. In your browser, nothing uploaded, which is rather the point when the concern is who's collecting your images.

The realistic playbook

For most people: turn off camera GPS tagging, clean metadata from photos you share as files, flip the AI-training toggles on platforms you use, and accept that anything posted publicly may be scraped — so choose what you post accordingly. For artists: add Glaze/Nightshade and in-file do-not-train signals to that list, and host your portfolio somewhere whose robots.txt you control. For everyone: don't pay for any service claiming it can retroactively remove your photos from training datasets. It can't.

Two of those steps are a browser tab away: check what a photo is carrying to see the GPS, device data, and AI tags it would hand a scraper, then add the do-not-train signal and your copyright before you publish it.

TL;DR

You can't technically prevent AI from scanning public photos — opt-out signals are voluntary, and pixels scrape regardless of metadata. What you can do: use the opt-outs anyway (they cover mainstream collectors), flip platform training toggles (contractually real), consider adversarial cloaking if you're an artist, and strip metadata from everything you share so that whatever does get scanned reveals your image but not your location, schedule, and devices. Control what's controllable; deny the rest its detail.

more in Privacy

see all →