Metadata-Version: 2.5
Name: audio-classifier-tool
Version: 1.3.0
Summary: Detect and cut target segments from audio — you define the classifier (ad removal is the flagship example), LLM-verified
Project-URL: Homepage, https://github.com/jon-fox/audio-classifier-tool
Project-URL: Repository, https://github.com/jon-fox/audio-classifier-tool
Author: Jon Fox
License-Expression: MIT
License-File: LICENSE
Keywords: ad-removal,ads,audio,classifier,podcast,whisper
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Multimedia :: Sound/Audio
Requires-Python: >=3.13
Requires-Dist: backoff>=2.2.1
Requires-Dist: boto3>=1.40.21
Requires-Dist: ctranslate2==4.8.2
Requires-Dist: faster-whisper==1.2.1
Requires-Dist: numpy>=2
Requires-Dist: nvidia-cublas-cu12; sys_platform == 'linux'
Requires-Dist: nvidia-cudnn-cu12==9.*; sys_platform == 'linux'
Requires-Dist: openai>=2
Requires-Dist: pydantic-ai-slim[openai]>=2
Requires-Dist: pydantic>=2.11.7
Requires-Dist: python-toon>=0.1.3
Requires-Dist: requests>=2.32.5
Requires-Dist: scikit-learn>=1.9.1
Requires-Dist: soundfile>=0.14.0
Provides-Extra: server
Requires-Dist: fastapi>=0.115; extra == 'server'
Requires-Dist: uvicorn>=0.30; extra == 'server'
Description-Content-Type: text/markdown

# AudioClassifier - Dynamic Classifier used for Audio Segment Identification and Removal

[![PyPI](https://img.shields.io/pypi/v/audio-classifier-tool)](https://pypi.org/project/audio-classifier-tool/) [![Python](https://img.shields.io/pypi/pyversions/audio-classifier-tool)](https://pypi.org/project/audio-classifier-tool/) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)

Runs locally by default: one audio file in, filtered audio out. Needs an OpenAI API key and `ffmpeg`.

## Quick Start

```bash
uv sync
export OPENAI_API_KEY=<your-key>

PAYLOAD='{"source": "My Show", "name": "Episode 1", "audio_url": "<mp3-url>"}' \
  uv run audioclassifier --detection examples/configs/ads.toon
```

Results land in `output/<source>/<name>/`: the original mp3, the cleaned `*_filtered.mp3`, a `segments.json` manifest (episode-absolute millisecond cut segments with confidence, plus the sha256/size/ETag identity of the exact copy analyzed — for clients that skip segments during playback instead of using the filtered file), and a `transcripts/` dir with what was transcribed and each LLM cut/keep decision (with reasoning). Or with Docker:

```bash
docker build -t audioclassifier-app .
docker run --gpus all \
  -e OPENAI_API_KEY=<your-key> \
  -e DETECTION_CONFIG=ads \
  -e PAYLOAD='{"source": "My Show", "name": "Episode 1", "audio_url": "<mp3-url>"}' \
  -v "$(pwd)/output:/app/output" \
  audioclassifier-app
```

Drop `--gpus all` to run on CPU (slower). The same container runs on any GPU box — RunPod, Modal, a gaming PC.

## As a Library

```bash
uv add audio-classifier-tool      # pip install audio-classifier-tool
```

```python
import audioclassifier

result = audioclassifier.process_audio(
    source="My Show",
    name="Episode 1",
    audio_url="<mp3-url>",
    detection="examples/configs/ads.toon",  # a .toon path or a name in ./configs
    detection_instructions="...",           # or pass the prompt directly
    detection_keywords=["use code", ...],   # and the keyword gate
)
print(result["output_path"], result["seconds_removed"])
```

Importing has no side effects; configure the `"audioclassifier"` logger to see progress.

For a real end-to-end run with live console output and a decision summary:

```bash
uv run python examples/process_audio.py "<mp3-url>"
```

## HTTP API (optional)

A FastAPI wrapper serves segment manifests to clients (e.g. podcast apps that skip ads client-side). Install the `server` extra and run it:

```bash
uv sync --extra server
API_TOKEN=<token> DETECTION_CONFIG=ads uv run uvicorn audioclassifier.server.app:app --host 0.0.0.0
```

Or `docker compose up server`. All requests need `Authorization: Bearer $API_TOKEN`.

- `POST /v1/episodes/analyze` with `{"episode_id", "source", "audio_url", "duration_sec"}` — returns `200` with the full manifest if this episode was already analyzed under the current classifier setup (idempotent; a changed `model_version` re-analyzes), else `202 {"status": "processing"}`. Jobs run one at a time (the pipeline is single-run per process).
- `GET /v1/episodes/{episode_id}/segments` — `{"status": "processing" | "failed" | "done", ...manifest fields when done}`.

Set `STORAGE=s3://bucket/prefix` to upload outputs; done responses then include `filtered_audio_url` so clients can fall back to the filtered MP3 when their copy of the episode doesn't match the analyzed one (`audio.duration_sec` / `audio.sha256`).

## How It Works

- Downloads the audio from the `PAYLOAD` JSON (`source`, `name`, `audio_url`, optional `description`)
- Transcribes with faster-whisper (GPU-accelerated, CPU works too)
- Detects target segments with an LLM, guided by your detection config (ad removal is the flagship example)
- Cuts them and re-assembles the audio

## Detection

The classifier is fully yours to define — nothing is bundled. A [TOON](https://github.com/toon-format/spec) config supplies the keywords and prompts: select one with `--detection <name|path>` or `DETECTION_CONFIG` (names resolve from `./configs/`), or pass the prompt and keywords directly (`detection_instructions`/`detection_keywords` in the API, `DETECTION_INSTRUCTIONS`/`DETECTION_KEYWORDS` env vars). Complete examples live in `examples/configs/` (ads, politics).

## Text Classifier (optional, opt-in)

A self-distilled local classifier can add a third detection signal alongside keywords and audio analysis. Every processed audio file writes LLM-labeled decisions under `output/**/transcripts/` — that corpus is the classifier's training data, and it grows with each run.

Enable with `USE_TEXT_CLASSIFIER=true`: the classifier's flagged ranges join the LLM prompt (advisory only — the LLM still decides), and it retrains automatically after each run from the **full accumulated history** (retraining is from scratch, in seconds, so nothing is ever forgotten). Refresh manually anytime:

```bash
uv run python examples/train_classifier.py          # or audioclassifier.train_text_classifier()
```

How it propagates: the durable memory is the decision files in `output/` — the model file (`local_models/text_classifier.joblib`) is a disposable cache rebuilt from them. Editing a decision file's `cut_ranges_seconds` after listening feeds your correction into the next training pass. In Docker, mount `local_models/` alongside `output/` to carry the model between containers (the data already survives via the `output/` mount).

## Parallelism

One audio file per `process_audio` call. Multiple processes are safe, even in the same directory — downloads and segment audio live in per-run temp dirs, and the classifier model is written atomically. Within one process, run audio files sequentially (Whisper models load once and are reused); concurrent runs in threads are supported only with one shared detection config. Don't feed the same audio to two processes at once — they'd write the same output files.

## Options

- `LLM_MODEL` — any [pydantic-ai model string](https://ai.pydantic.dev/models/) (default `openai:gpt-5.6`; e.g. `openai:gpt-5-nano` for cheapest, `anthropic:claude-sonnet-4-6`, `ollama:qwen3` — non-OpenAI providers may need their extra installed)
- `DISCORD_ALERTS=true DISCORD_WEBHOOK_URL=<url>` — processing alerts in Discord
- `storage="s3://bucket/prefix"` (API) or `"storage"` in `PAYLOAD` — upload the run's outputs (audio + transcripts + decisions) to S3 under `<prefix>/<source>/<name>/`, using ambient AWS credentials

## License

[MIT](LICENSE)
