Metadata-Version: 2.4
Name: topicscout
Version: 0.2.0
Summary: Find the GitHub repos on a topic, and only the new ones next time. No dependencies, token optional.
Author: Adecubed
License-Expression: MIT
Project-URL: Homepage, https://github.com/adecubed/topicscout
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: mcp
Requires-Dist: mcp>=1.2; extra == "mcp"
Provides-Extra: keywords
Requires-Dist: github; extra == "keywords"
Requires-Dist: scraper; extra == "keywords"
Requires-Dist: search; extra == "keywords"
Requires-Dist: discovery; extra == "keywords"
Requires-Dist: topics; extra == "keywords"
Requires-Dist: mcp; extra == "keywords"
Dynamic: license-file

# topicscout

[![ci](https://github.com/adecubed/topicscout/actions/workflows/ci.yml/badge.svg)](https://github.com/adecubed/topicscout/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/topicscout)](https://pypi.org/project/topicscout/)

Find the GitHub repositories on a topic, and next time only the new ones.

```bash
pip install topicscout        # or run it without installing: uvx topicscout run github-scrapers
topicscout run github-scrapers
```

```
19 found, 17 kept, 17 new -> scout-github-scrapers/new.md
```

`new.md` is a table of what appeared since the last run, most stars first: stars,
last push, language, license, a guess at the interface from the README (MCP, pip,
npm, REST, Docker...) and the description. `candidates.json` keeps everything seen,
so a weekly run is a short list, not the same 300 repos again.

No dependencies (the MCP server is an optional extra), Python 3.11+. `GITHUB_TOKEN` is optional: without it the searches
are paced to stay under GitHub's 10 per minute and READMEs are capped at 40 per run.

## Why

GitHub search is good at "repos with this topic" and bad at "what's new since I last
looked, minus the forks, the awesome-lists, the abandoned ones and the repos that
only tagged themselves". topicscout runs a fixed set of searches, applies your
filters, remembers what it has seen and what you already have.

It was the discovery step of [adebench](https://github.com/adecubed/adebench), a
benchmark for AI memory systems, which uses it to find new memories to measure.

## Your own topic

Copy a built-in profile (they are in [`topicscout/profiles/`](topicscout/profiles/)),
change the queries and the words a hit must contain, and run it:

```bash
topicscout run my-topic.json
```

Run it again next week: `new.md` lists only what appeared in between. Put the repos
you already know or rejected in `scout-my-topic/known.txt`.

## From a script or an agent

`--json` prints the result on stdout, progress stays on stderr:

```bash
topicscout run github-scrapers --json -q
```

```json
{"profile": "github-scrapers", "found": 19, "kept": 17, "new": [{"full_name": "yusufkaraaslan/Skill_Seekers", "stars": 15052, "interface": ["cli", "pip", "action"], "...": "..."}, "..."], "stopped": null, "out": "scout-github-scrapers"}
```

It is also a Python library (`from topicscout.core import Profile, run`) and an MCP server.

## As an MCP server

```bash
pip install "topicscout[mcp]"
claude mcp add topicscout -- topicscout mcp      # Claude Code; any MCP client runs `topicscout mcp`
```

Or in a client's JSON config: `{"command": "topicscout", "args": ["mcp"], "env": {"GITHUB_TOKEN": "..."}}`.

| tool | what it does | network |
|---|---|---|
| `list_profiles` | built-in profiles and the state folder | no |
| `scout_run` | search GitHub with a profile, return what is new | yes, slow without a token |
| `scout_candidates` | what past runs saw: filter by status (`new`, `seen`, `known`), text, stars | no |
| `scout_mark_known` | add a repo to `known.txt` with a note, so it never comes back as new | no |
| `doctor` | token and remaining GitHub quota | yes |

The state of each profile lives in `~/.topicscout/<profile>` (or `$TOPICSCOUT_HOME`);
every tool also takes `out`, the same folder the CLI's `--out` takes, so the agent and
the command line can share one state.

## Commands

```bash
topicscout run PROFILE [--out DIR] [--min-stars N] [--days N] [--readme-max N]
                       [--known-from GLOB ...] [--json] [-q]
topicscout profiles    # built-in profiles
topicscout doctor      # token and remaining GitHub quota
topicscout mcp         # MCP server on stdio (needs topicscout[mcp])
```

- `PROFILE` is a built-in name or a path to your own JSON profile.
- `--out` is the state folder (default `scout-<profile>`): `candidates.json`,
  `new.md`, and your `known.txt`.
- `--json` prints the summary and the new repos as JSON on stdout; progress and
  errors always go to stderr, so scripts and agents can read stdout as is.

## What counts as known

Repos you already have are marked `known` and never listed as new:

- `known.txt` in the state folder: one `owner/repo` per line, `#` comments. Put
  rejected repos here too.
- `--known-from "adapters/*.py"`: the GitHub URLs in the first 20 lines of those
  files. adebench points it at its adapters, so a memory with an adapter drops out.

## Profiles

Built in: `ai-memory` (memory systems for AI agents and LLMs) and `github-scrapers`
(tools that scrape, crawl or mine GitHub). A profile is a JSON file:

```json
{
  "name": "ai-memory",
  "description": "Memory systems for AI agents and LLMs",
  "queries": ["topic:agent-memory", "\"llm memory\" in:name,description"],
  "exclude": "(?i)awesome|curated list|\\bpapers\\b",
  "require": [
    {"pattern": "(?i)memor|\\bmem\\b", "own_words": true},
    {"pattern": "(?i)\\b(agents?|llms?|ai|mcp)\\b"}
  ],
  "interfaces": [["mcp", "(?i)\\bmcp\\b"], ["pip", "(?i)\\bpip install\\b"]],
  "min_stars": 10,
  "days": 90
}
```

- `queries`: [GitHub repository search](https://docs.github.com/en/search-github/searching-on-github/searching-for-repositories)
  queries. Each one always gets `fork:false archived:false stars:>=min_stars
  pushed:>=today-days`; up to 300 results per query.
- `exclude`: a regex on name, description and topics; a match drops the repo.
- `require`: every pattern must match. With `own_words` it must be in the name or
  description, not only in a topic: a database that tagged itself `agent-memory`
  is not a memory. Hyphens and underscores count as spaces.
- `interfaces`: `[label, regex]` pairs tried on the README of each new repo.
  Leave it out and no README is fetched.
- `require_language` (default true) drops repos with no language, which are
  usually lists and docs.

Patterns are Python regexes; add `(?i)` for case-insensitive.

## Rate limits

On a 403 or 429 it waits for the reset if that is under 70 seconds; otherwise it
stops, saves what it found and says so in `new.md` and on stderr. The next run
picks up the READMEs it did not get to.

## How it compares

- [ghcrawl](https://github.com/pwrdrvr/ghcrawl) goes deep into one repository
  (issues and PRs, embeddings, clusters); topicscout goes wide across GitHub for a
  topic.
- [top-github-scraper](https://github.com/khuyentran1401/top-github-scraper)
  lists the top repos for a keyword, once; topicscout adds filters, profiles and
  memory across runs.

## License

MIT
