Metadata-Version: 2.5
Name: imgtrail
Version: 0.1.2
Summary: Find out where else on the web your own photos show up
Project-URL: Homepage, https://github.com/Endika/imgtrail
Project-URL: Repository, https://github.com/Endika/imgtrail
Project-URL: Issues, https://github.com/Endika/imgtrail/issues
Project-URL: Changelog, https://github.com/Endika/imgtrail/blob/main/CHANGELOG.md
Author-email: Endika Iglesias <endika2@gmail.com>
License-Expression: MIT
License-File: LICENSE
Keywords: instagram,osint,phash,privacy,reverse-image-search
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: End Users/Desktop
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Multimedia :: Graphics
Classifier: Topic :: Security
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: httpx>=0.27
Requires-Dist: imagehash>=4.3
Requires-Dist: pillow>=10
Requires-Dist: rich>=13
Description-Content-Type: text/markdown

# imgtrail

[![PyPI](https://img.shields.io/pypi/v/imgtrail)](https://pypi.org/project/imgtrail/)
[![Python](https://img.shields.io/pypi/pyversions/imgtrail)](https://pypi.org/project/imgtrail/)
[![CI](https://github.com/Endika/imgtrail/actions/workflows/ci.yml/badge.svg)](https://github.com/Endika/imgtrail/actions/workflows/ci.yml)
[![Licence](https://img.shields.io/pypi/l/imgtrail)](LICENSE)
[![Checked with mypy](https://img.shields.io/badge/mypy-strict-2a6db2)](https://mypy-lang.org/)
[![Ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json)](https://github.com/astral-sh/ruff)

Find out where else on the web your own photos show up.

Point it at your Instagram data export. It hashes every photo, collapses the near-duplicates
so you never pay to search the same picture twice, runs each unique one through reverse image
search, and then **downloads every candidate and compares it against your original** before
putting it in the report. What you get back is a list you can trust, not a pile of URLs.

```
imgtrail scan ~/Downloads/instagram-export.zip --dry-run
imgtrail scan ~/Downloads/instagram-export.zip
imgtrail report --open
```

## What it finds, and what it doesn't

It searches Google's index, so it finds your photos on **blogs, news sites, Pinterest, Tumblr,
forums, scraper mirrors and shops that lifted your pictures**.

It will **not** find a repost on another Instagram account. Instagram blocks crawling of post
images, so they aren't in anyone's index — the only way such a repost surfaces here is
indirectly, via one of the many "Instagram viewer" mirror sites that *are* indexed. Telegram,
WhatsApp, TikTok, Facebook and private accounts are invisible to it too. If your question is
"is someone reposting me inside Instagram", this is the wrong tool and there isn't a good one.

## Install

```
pip install imgtrail
```

## Getting your photos

Instagram → Settings → Accounts Centre → Your information and permissions → **Download your
information**. Ask for JSON, high quality. You'll get a ZIP; hand it straight to `imgtrail scan`.
No scraping, nothing against the terms of service, no rate limits.

A plain folder of images works just as well.

## Getting an API key

imgtrail uses Google Cloud Vision's `WEB_DETECTION`. Create a project at
[console.cloud.google.com](https://console.cloud.google.com), enable the **Cloud Vision API**,
then Credentials → Create credentials → API key.

```
export IMGTRAIL_API_KEY=AIza...
```

**The first 1,000 images each month are free**, then $3.50 per 1,000. A typical profile costs
nothing. Run `--dry-run` first and it will tell you exactly how many searches it would make and
what they would cost before spending anything.

## How the verification works

Reverse image search returns a lot of near-misses. For every candidate, imgtrail downloads the
image and compares perceptual hashes against your original:

| Hamming distance | Verdict | Meaning |
|---|---|---|
| ≤ 8 | `confirmed` | The same image, possibly recompressed |
| ≤ 16 | `likely` | Cropped, filtered or heavily edited |
| > 16 | `rejected` | Not your photo |

Only `confirmed` and `likely` reach the report. `visuallySimilarImages` is dropped entirely —
it means "semantically alike", not "this is your photo", and it drowns the report in noise.

## Commands

```
imgtrail scan SOURCE          index, dedupe, search and verify — resumable
  --dry-run                   count the searches and their cost, call nothing
  --limit N                   search at most N unique photos
  --threshold N               pHash distance for "same photo" (default 6)
  --ignore-domain DOMAIN      exclude a domain from results (repeatable)
  --no-verify                 skip the download-and-compare pass
imgtrail report --open        build the HTML report and open it
imgtrail status               what's in the database so far
```

State lives in `./imgtrail-data`. Everything is idempotent: re-running `scan` searches only
what it hasn't searched before, so an interrupted run costs nothing to resume.

## Privacy

Your photos are sent to Google Cloud Vision, and nowhere else. Nothing is uploaded to any
server of mine — there isn't one. The database, the extracted export and the report all stay
on your machine.

## Architecture

Ports and adapters, sized to the problem: the rules sit in the middle and know nothing
about Google, SQLite or HTTP, so swapping a search backend touches exactly one file.

```
domain.py     fingerprints, grouping, verdicts, what counts as "your own platform"
              — pure; no I/O, no SQL, no network
ports.py      the boundaries: PhotoSource, ImageLoader, SearchEngine, ImageFetcher,
              PhotoRepository, MatchRepository, ReportWriter
services.py   the use cases: index, plan, search, verify, report
adapters/     the details: sqlite_repository, vision, http_fetcher, local_files, html_report
cli.py        the composition root — the one module that knows every layer
```

Adding TinEye or Yandex means writing one `SearchEngine` and wiring it in `cli.py`. Nothing
in `domain.py` or `services.py` changes.

## Development

```bash
uv sync --all-groups
uv run pytest             # 85 tests, no network, no mocks
uv run ruff check .
uv run ruff format .
uv run mypy               # strict, and it passes on the tests too
```

The test doubles are real implementations, not mocks: an in-memory `DictPhotoSource`, a
`FakeSearchEngine` that records what it was asked, and — where the wire itself is what needs
testing — a real local HTTP server speaking Vision's JSON.

## Licence

MIT
