Metadata-Version: 2.5
Name: subxt
Version: 1.0.1.post1
Summary: End-to-end subtitle generation, translation, and multi-format export CLI
Author: subxt contributors
License-Expression: MIT
License-File: LICENSE
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: <3.14,>=3.12
Requires-Dist: gemini-srt-translator>=3.8.0
Requires-Dist: pydantic-settings>=2.2.0
Requires-Dist: pydantic>=2.7.0
Requires-Dist: pysubs2>=1.7.0
Requires-Dist: python-dotenv>=1.0.0
Requires-Dist: rich>=13.7.0
Requires-Dist: typer>=0.12.0
Requires-Dist: whisperx>=3.1.0
Requires-Dist: yt-dlp>=2024.4.0
Provides-Extra: dev
Requires-Dist: black>=24.0.0; extra == 'dev'
Requires-Dist: pytest-mock>=3.12.0; extra == 'dev'
Requires-Dist: pytest>=8.0.0; extra == 'dev'
Requires-Dist: ruff>=0.4.0; extra == 'dev'
Description-Content-Type: text/markdown

# subxt

**`subxt`** is a production-quality CLI tool that automates end-to-end subtitle creation: optionally download source media with `yt-dlp`, transcribe and align speech with `whisperX`, translate transcripts into target languages with `gemini-srt-translator`, and export results into any subtitle and caption formats you need.

Every stage works **standalone** (its own subcommand) and **chained** via `subxt run`.

---

## Features

- **Optional Media Download:** Seamlessly fetch audio/video from 1,000+ sites with `yt-dlp`.
- **Word-Level Transcription & Diarization:** GPU/CPU transcription and forced alignment with `whisperX`, plus optional speaker diarization.
- **AI-Powered Subtitle Translation:** Fast, context-aware SRT translation using Google Gemini via `gemini-srt-translator`.
- **Bilingual Subtitles:** Output combined original + translated subtitles in `.srt` or styled `.ass`.
- **Comprehensive Format Support:** Export to SubRip (`.srt`), WebVTT (`.vtt`), JSON (`.json`), LRC (`.lrc` with enhanced word timings), Plain text (`.txt`), CSV (`.csv`), SubViewer (`.sbv`), and legacy formats (`.ass`, `.ssa`, `.sub`, `.mpl2`, `.tmp`, `.ttml`, `.smi`).
- **Resumable & Safe:** Track progress in `manifest.json`; skip already-completed expensive stages with `--resume`.
- **Dry-run Planning:** Preview resolved plans and output paths with `--dry-run` before spending API quota or GPU time.

---

## Installation

### Prerequisites

- **Python:** `>=3.12, <3.14`
- **FFmpeg:** Required on `PATH` for audio decoding and extraction.
  ```bash
  # Linux (Ubuntu/Debian)
  sudo apt update && sudo apt install ffmpeg
  # macOS
  brew install ffmpeg
  # Windows
  winget install Gyan.FFmpeg
  ```

### Basic Installation

For regular users, install `subxt` using `pip`:

```bash
pip install subxt
```

### Advanced Installation

For developers and contributors wanting to modify or develop `subxt` locally:

```bash
git clone https://github.com/Iceyu29/subxt.git
cd subxt
pip install -e .
```

### CUDA Installation

To use WhisperX with GPU acceleration, install the CUDA toolkit 12.8 before WhisperX. Skip this step if using only the CPU.

- For **Linux** users, install the CUDA toolkit 12.8 following this guide: [CUDA Installation Guide for Linux](https://docs.nvidia.com/cuda/cuda-installation-guide-linux/).
- For **Windows** users, download and install the CUDA toolkit 12.8: [CUDA Downloads](https://developer.nvidia.com/cuda-downloads).

By default, `subxt` auto-detects CUDA availability. If no GPU is found, it automatically falls back to CPU with `int8` computation.

---

## Configuration & Precedence

`subxt` resolves configuration with strict precedence:
```
Explicit CLI Flag > Environment Variable / .env > Configuration File (subxt.toml) > Built-in Default
```

### Setting Up API Keys & Tokens

`subxt` requires a Google Gemini API key for subtitle translation and an optional Hugging Face token if you enable speaker diarization (`--diarize`).

#### 1. Google Gemini API Key (`GEMINI_API_KEY`)
Required for translating subtitles (`subxt translate` and `subxt run -t <lang>`):
1. Visit [Google AI Studio](https://aistudio.google.com/).
2. Sign in with your Google account.
3. Click **"Get API key"** -> **"Create API key"**.
4. *(Optional)* Create a secondary key (`GEMINI_API_KEY2`) for automatic quota failover when hitting free-tier limits.

#### 2. Hugging Face Access Token (`HF_TOKEN`)
Required **only** if you enable speaker diarization (`--diarize`):
1. Sign up or log in at [Hugging Face](https://huggingface.co/).
2. Go to [Access Tokens Settings](https://huggingface.co/settings/tokens) and click **"Create new token"** (a **Read** token is sufficient).
3. **Accept Model User Agreements:** Diarization models require accepting the author terms while logged into Hugging Face:
   - Accept terms at: [pyannote/speaker-diarization-community-1](https://huggingface.co/pyannote/speaker-diarization-community-1) (or [pyannote/speaker-diarization-3.1](https://huggingface.co/pyannote/speaker-diarization-3.1))
   - Accept terms at: [pyannote/segmentation-3.0](https://huggingface.co/pyannote/segmentation-3.0)

### How to Configure Secrets

You can supply your credentials in three ways (highest precedence wins):

- **Option A: `.env` file (Recommended)**
  Create a `.env` file in your working directory (or project root):
  ```env
  GEMINI_API_KEY="AIzaSy..."
  GEMINI_API_KEY2="AIzaSy..."  # Optional secondary key for quota failover
  HF_TOKEN="hf_..."           # Optional for speaker diarization
  ```

- **Option B: Export as Environment Variables**
  In your terminal session (or add to `~/.bashrc` / `~/.zshrc`):
  ```bash
  export GEMINI_API_KEY="AIzaSy..."
  export HF_TOKEN="hf_..."
  ```

- **Option C: Pass via CLI Flags**
  Pass keys on-demand in the command line:
  ```bash
  subxt run "video.mp4" -t fr --gemini-api-key "AIzaSy..." --diarize --hf-token "hf_..."
  ```

### `subxt.toml` Configuration File

Create a `subxt.toml` in your working directory to set default preferences:

```toml
whisper_model = "large-v2"
gemini_model = "gemini-3.5-flash"
export_formats = ["srt", "vtt", "json"]
device = "cuda"
bilingual = true
bilingual_format = "srt"
output_dir_template = "./subxt-output/{slug}"
```

---

## Quickstart

### 1. Full Pipeline (Download, Transcribe, Translate, Export)

```bash
export GEMINI_API_KEY="your-gemini-key"

# Download video, transcribe with whisperX, translate to French, and export to SRT + VTT
subxt run "https://www.youtube.com/watch?v=dQw4w9WgXcQ" -t fr -f srt,vtt --bilingual
```

### 2. Local Media Pipeline (Transcription-Only)

```bash
subxt run ./lecture.mp4 --no-translate -f srt,vtt,json,txt
```

### 3. Individual Subcommands

```bash
# 1. Download audio only
subxt download "https://www.youtube.com/watch?v=dQw4w9WgXcQ" --audio-only -o ./downloads

# 2. Transcribe local audio file
subxt transcribe ./downloads/audio.wav --whisper-model large-v2 --device cuda

# 3. Translate an existing SRT file
subxt translate ./downloads/audio.source.srt -t es --gemini-model gemini-3.5-flash

# 4. Export subtitles to multiple formats
subxt export ./downloads/audio.source.srt --translated ./downloads/audio.es.srt -f vtt,lrc,json,csv
```

---

## Supported Export Formats

| Format | Extension | Engine / Method | Description |
|---|---|---|---|
| **SubRip** | `.srt` | pysubs2 | Standard subtitle format, always generated |
| **WebVTT** | `.vtt` | pysubs2 | Modern HTML5 web captions |
| **JSON** | `.json` | Custom Schema | Rich payload: original, translation, speaker, and word timings |
| **Plain Text** | `.txt` | Custom | Clean transcript (`--txt-style lines` or `paragraph`) |
| **LRC** | `.lrc` | Custom | Lyrics timing; enhanced word-level `<mm:ss.xx>` when word timings exist |
| **CSV** | `.csv` | Custom | Spreadsheet for proofreading (`index,start,end,original,translated,speaker`) |
| **Bilingual SRT** | `.srt` | Custom merge | Two lines per cue (`--bilingual-order original_first\|translated_first`) |
| **Bilingual ASS** | `.ass` | pysubs2 + Custom | Two-language styled subtitles with distinct visual formatting |
| **Advanced SSA** | `.ass` | pysubs2 | Styled subtitle format with typography support |
| **SubStation Alpha** | `.ssa` | pysubs2 | Legacy SSA format |
| **MicroDVD** | `.sub` | pysubs2 | Frame-based legacy subtitle format (default 25 fps) |
| **MPL2** | `.mpl2` | pysubs2 | Time-based subtitle format |
| **TMP** | `.tmp` | pysubs2 | Legacy start-time only format |
| **TTML** | `.ttml` | pysubs2 | Timed Text Markup Language for broadcast/streaming QC |
| **SAMI** | `.smi`, `.sami` | Custom compliant | Synchronized Accessible Media Interchange format |
| **SubViewer** | `.sbv` | Custom | YouTube legacy caption format (`h:mm:ss.mmm,h:mm:ss.mmm`) |

---

## CLI Reference

### `subxt run`

```
Usage: subxt run [OPTIONS] INPUT

  Run the complete end-to-end pipeline: download, transcribe, translate, export.
  All options from individual stages (download, transcribe, translate, export) can
  be passed directly to `subxt run` (e.g. `--audio-only`, `--whisper-model`,
  `--gemini-model`, `--bilingual`, `--resume`).

Arguments:
  INPUT                            URL or local file path. [required]

Options:
  -t, --target-language TEXT       Language(s) to translate into (comma-separated, e.g. 'fr,es').
  --skip-download                  Treat input as local file even if it looks like a URL.
  --no-translate                   Skip translation step (transcription only).
  -f, --export-formats TEXT        Comma-separated export formats (default: 'srt').
  --bilingual / --no-bilingual     Emit combined bilingual subtitle file.
  --bilingual-format [srt|ass]     Format for bilingual file (default: srt).
  --bilingual-order [original_first|translated_first]
  -o, --output-dir PATH            Output directory (default: ./subxt-output/<slug>).
  --resume / --no-resume           Skip steps whose outputs already exist in manifest.
  --dry-run                        Print resolved execution plan and exit.
  --config PATH                    Path to TOML config file.
  -v, --verbose                    Enable verbose debug output.
  -q, --quiet                      Suppress all output except errors.
  --log-file PATH                  Log output to specified file.
```

### `subxt download`

```
Usage: subxt download [OPTIONS] URL_OR_FILE

Options:
  --audio-only                     Extract audio only.
  --format TEXT                    yt-dlp format selector (default: bestvideo*+bestaudio/best for video, bestaudio/best for audio-only).
  --cookies-from-browser TEXT      Browser name to load cookies from.
  -o, --output-dir PATH            Output directory.
```

### `subxt transcribe`

```
Usage: subxt transcribe [OPTIONS] MEDIA

Options:
  --whisper-model TEXT             Whisper model name (default: large-v2).
  --source-language TEXT           ISO language code, or auto-detect if omitted.
  --device [auto|cuda|cpu]         Execution device.
  --compute-type TEXT              Quantization/compute type (float16/int8/float32).
  --whisper-batch-size INTEGER     Inference batch size (default: 16).
  --whisper-chunk-size INTEGER     WhisperX chunk size (default: 5).
  --no-align                       Disable forced word-level alignment.
  --align-model TEXT               Override wav2vec2 alignment model.
  --diarize                        Enable speaker diarization.
  --hf-token TEXT                  Hugging Face token for diarization.
  --min-speakers INTEGER           Minimum speaker hint.
  --max-speakers INTEGER           Maximum speaker hint.
  -o, --output-dir PATH            Output directory.
```

### `subxt translate`

```
Usage: subxt translate [OPTIONS] SRT_FILE

Options:
  -t, --target-language TEXT       Target language (e.g. French, Spanish, ja) [required].
  --gemini-api-key TEXT            Google Gemini API key.
  --gemini-model TEXT              Model name (default: gemini-3.5-flash).
  --translate-batch-size INTEGER   Cue batch size (default: 1000).
  --description TEXT               Contextual description for translation.
  --start-line INTEGER             Resume translation from line number.
  -o, --output-path PATH           Output file path.
```

### `subxt export`

```
Usage: subxt export [OPTIONS] SOURCE_SRT

Options:
  --translated PATH                Path to translated SRT file.
  -f, --formats TEXT               Comma-separated list of formats to export.
  --raw-json PATH                  Path to raw whisperX JSON for word timings.
  --bilingual / --no-bilingual     Export combined bilingual subtitle file.
  --bilingual-format [srt|ass]     Bilingual format.
  --txt-style [lines|paragraph]    Plain text formatting style.
  -o, --output-dir PATH            Output directory.
```

### `subxt list-models`

```
Usage: subxt list-models [OPTIONS]

  List available Gemini models for translation via gemini-srt-translator.
```

---

## Exit Codes

`subxt` returns specific non-zero exit codes to enable clean scripting:
- `0`: Success
- `1`: Unhandled error
- `2`: Configuration or CLI validation error
- `10`: Media download error
- `20`: Transcription or alignment error
- `30`: Translation or cue validation error
- `40`: Export or formatting error

---

## Credits & Acknowledgements

`subxt` is built on top of and powered by these open-source projects:

- **[whisperX](https://github.com/m-bain/whisperX)** by [Max Bain](https://github.com/m-bain) et al. (BSD-2-Clause) — Fast batched speech transcription, phoneme-level forced alignment, and speaker diarization.
- **[gemini-srt-translator](https://github.com/MaKTaiL/gemini-srt-translator)** by [MaKTaiL](https://github.com/MaKTaiL) (MIT) — Context-aware SRT subtitle translation powered by Google Gemini.
- **[yt-dlp](https://github.com/yt-dlp/yt-dlp)** (Unlicense) — Command-line audio and video download engine supporting 1,000+ sites.
- **[pysubs2](https://github.com/tkarabela/pysubs2)** by [Tomas Karabela](https://github.com/tkarabela) (MIT) — Subtitle manipulation for SubRip, WebVTT, ASS/SSA, MicroDVD, TTML, etc.
- **[pyannote-audio](https://github.com/pyannote/pyannote-audio)** by [Hervé Bredin](https://github.com/hbredin) et al. (MIT) — State-of-the-art neural speaker diarization pipelines.
- **[faster-whisper](https://github.com/SYSTRAN/faster-whisper)** & **[CTranslate2](https://github.com/OpenNMT/CTranslate2)** (MIT) — Fast C++ inference engine for Whisper models on CPU and GPU.
- **[Typer](https://github.com/fastapi/typer)** by [Sebastián Ramírez](https://github.com/tiangolo) (MIT) — CLI library powering type-hinted commands and help output.
- **[Rich](https://github.com/Textualize/rich)** by [Will McGugan](https://github.com/willmcgugan) (MIT) — Terminal formatting, tables, and progress bars.
- **[FFmpeg](https://ffmpeg.org/)** (LGPL/GPL) — Multimedia framework for audio extraction and conversion.

---

## License

MIT License. See [LICENSE](LICENSE) for details.
