Metadata-Version: 2.5
Name: airelays
Version: 0.14.0
Summary: AIRelays: an independent OpenAI-compatible local relay for single-user subscription-backed access.
Project-URL: GitHub, https://github.com/lpalbou/AIRelays
Project-URL: Documentation, https://www.lpalbou.info/AIRelays/
Project-URL: Changelog, https://github.com/lpalbou/AIRelays/blob/main/CHANGELOG.md
Author: Laurent-Philippe Albou
License: MIT
License-File: LICENSE
Keywords: ai,gateway,openai-compatible,relay,subscription
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Framework :: FastAPI
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Requires-Python: >=3.11
Requires-Dist: fastapi>=0.115.0
Requires-Dist: httpx>=0.28.0
Requires-Dist: keyring>=25.5.0
Requires-Dist: python-multipart>=0.0.20
Requires-Dist: uvicorn>=0.35.0
Provides-Extra: dev
Requires-Dist: build>=1.2.0; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.24.0; extra == 'dev'
Requires-Dist: pytest>=8.3.0; extra == 'dev'
Requires-Dist: twine>=6.1.0; extra == 'dev'
Provides-Extra: docs
Requires-Dist: mkdocs-material>=9.0.0; extra == 'docs'
Requires-Dist: mkdocs>=1.6.0; extra == 'docs'
Requires-Dist: pymdown-extensions>=10.0; extra == 'docs'
Description-Content-Type: text/markdown

# AIRelays

`AIRelays` is a local OpenAI-compatible HTTP server with provider-scoped runtimes.

- The default runtime uses an AIRelays-owned ChatGPT subscription login.
- An optional Claude runtime uses the local `claude` CLI and its existing subscription auth state.
- AIRelays protects the relay with its own bearer token by default.
- Traffic is logged to JSONL files with automatic rotation and retention (7 days, 1 GiB total by default).

## Independence And Intended Use

- AIRelays is an independent third-party project. It is not affiliated with, endorsed by, or sponsored by any provider.
- Provider and product names are used only to describe compatibility targets and upstream behavior.
- AIRelays is designed for a single user running a local relay for personal convenience.
- AIRelays is not presented as a shared, pooled, multi-user, or resale service.
- The Claude runtime is local-only and not presented as a sanctioned provider integration path.
- You are responsible for complying with each provider's terms. Both providers currently frame subscription access around ordinary, individual use by the account holder; the moment anyone else's requests flow through your relay, you are outside that.

See [DISCLAIMER.md](DISCLAIMER.md) — it links the official Anthropic and OpenAI terms and policy pages to review.

## Install

AIRelays ships two ways; both drive the same relay and share the same
config (`~/.config/airelays`) and data (`~/.airelays`).

### CLI / server install (PyPI)

For headless machines, servers, or terminal-first workflows:

```bash
python -m pip install airelays
```

Or from a source checkout:

```bash
python -m pip install .
```

### Desktop app (GUI + system tray)

A cross-platform tray app (macOS, Windows, Linux) lives under
[desktop/](desktop/README.md): a dashboard with relay start/stop, auth and
network modes, OpenAI and Claude sign-in/sign-out, per-account usage bars,
a model list with copy-ready ids, live traffic, and diagnostics. The tray
icon shows connection state and blinks on request activity; the app can
start at login, starts the relay when it opens, and restarts a crashed
relay automatically. Installers (DMG, NSIS, AppImage, deb) build from
`.github/workflows/desktop.yml`; locally:

```bash
cd desktop
./scripts/bundle_runtime.sh
npm install && npm run build
```

An earlier native macOS status-bar app remains available under
[macos/AIRelaysMenuBar](macos/AIRelaysMenuBar/README.md):

```bash
swift build --package-path macos/AIRelaysMenuBar
swift run --package-path macos/AIRelaysMenuBar AIRelaysMenuBar
```

## Quick Start

OpenAI runtime:

```bash
airelays init
airelays login
airelays doctor
airelays serve --port 8080
```

Headless / server install (SSH, no browser):

```bash
airelays init
airelays login --device
airelays doctor
airelays serve --port 8080
```

Device-code login prints a short code you approve from a browser on any
other device (laptop, phone). On SSH sessions and displayless Linux,
`airelays login` selects it automatically. Do not paste the browser-flow
URL into a browser on another computer: its sign-in redirect only works on
the machine running the relay (see
[docs/troubleshooting.md](docs/troubleshooting.md) for the SSH-tunnel
alternative).

OpenAI runtime in open local relay mode:

```bash
airelays init --no-auth
airelays login
airelays serve --no-auth --port 8080
```

This disables only the AIRelays client-token gate. It does not bypass the upstream ChatGPT login.

Claude runtime:

```bash
airelays init
claude auth login --claudeai
airelays serve --port 8080
```

Claude runtime in headless environments:

```bash
# on any machine WITH a browser:
claude setup-token          # prints a long-lived token

# on the server:
airelays init
airelays claude set-token   # paste the token; stored 0600, survives restarts
airelays serve --port 8080
```

`airelays claude set-token` stores the token in `~/.airelays/claude-token`
and passes it to the local `claude` CLI automatically — unlike a shell
`export`, it keeps working under systemd, launchd, and docker. Exporting
`CLAUDE_CODE_OAUTH_TOKEN` still works as a fallback. `airelays claude
logout` signs Claude out completely: it removes the stored token and runs
`claude auth logout` (which signs out every tool using the `claude` CLI on
that machine).

When the Claude runtime is enabled, AIRelays keeps the same auth behavior as the rest of the relay. The default protected mode requires the AIRelays bearer token; `--no-auth` starts an open local relay. Claude remains restricted to loopback binding.

## Basic Verification

Run setup and upstream probes before starting the server:

```bash
airelays doctor
```

List the model ids the running relay accepts (grouped by provider):

```bash
airelays models
```

Use `airelays doctor --skip-response` to skip the tiny `/responses` smoke request.

Public health:

```bash
curl http://127.0.0.1:8080/healthz
```

Protected relay status:

```bash
curl http://127.0.0.1:8080/v1/relay/status \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN'
```

OpenAI model listing:

```bash
curl http://127.0.0.1:8080/v1/models \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN'
```

OpenAI text request:

```bash
curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN' \
  -H 'content-type: application/json' \
  -d '{
    "model": "gpt-5.5",
    "messages": [{"role": "user", "content": "Reply with exactly: OPENAI AIRelays OK"}]
  }'
```

Claude text request:

```bash
curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN' \
  -H 'content-type: application/json' \
  -d '{
    "model": "claude:sonnet",
    "messages": [{"role": "user", "content": "Reply with exactly: CLAUDE AIRelays OK"}]
  }'
```

## Multiple OpenAI Accounts

One person can enroll several of their own OpenAI subscriptions and let the
relay balance across them. Signing in again with a different account adds it
(the previous sign-in is kept, never overwritten):

```bash
airelays login            # first account
airelays login            # second account — added alongside the first
airelays accounts         # list accounts and the commands to manage them
airelays logout perso@gmail.com          # sign one account out
airelays accounts order work@company.com perso@gmail.com   # change priority
```

`airelays accounts` is the hub: it lists your accounts in balancing order and
prints the exact commands to add, sign out, or reorder them. In the desktop
app, each account row has a sign-out button and the "Add account" button
offers both browser and code (headless) sign-in.

By default AIRelays balances requests by remaining capacity
(`balance = "balanced"`): the account with the most unused weekly quota
serves next, so consumption equalizes as a percentage of each plan's own
capacity — a small Plus plan and a large Enterprise plan deplete
proportionally instead of the small plan draining many times faster.
Which usage windows a plan reports is upstream policy (some plans report
only a weekly window), so the relay identifies windows by duration and
balances on each account's longest one. The relay probes each account's
usage at launch and refreshes it in the background; an account at a
usage limit is benched until that window resets and rejoins rotation
automatically. Alternatives:
`balance = "round_robin"` sends strictly equal request counts, and
`balance = "ordered"` drains the first account before touching the next.
Failed-over requests are logged with the serving account, and
`/v1/subscription/status?all_accounts=true` reports usage per account.

Multiple accounts exist so one user can use their own subscriptions from
one relay; it is not a mechanism for sharing or pooling access between
people (see [DISCLAIMER.md](DISCLAIMER.md)).

## Relay Token

Show the current token:

```bash
airelays token show
```

Rotate the current token:

```bash
airelays token rotate
```

Use the relay token as the client credential when you point an OpenAI-compatible SDK at AIRelays.

## Provider Routing

- models starting with `claude:` or `claude-` use the Claude runtime when it is enabled
- other model ids use the OpenAI runtime when it is enabled
- OpenAI model discovery queries the upstream catalog using the installed Codex client version, with a bundled version floor for standalone installations. Optional `[providers.openai] extra_models` entries extend the catalog; other unlisted ids are rejected locally when the catalog is available.
- Claude model discovery queries the installed CLI and exposes both aliases and concrete model ids. The Models tab and `airelays models` show each reported alias resolution, such as `claude:fable` resolving to `claude-fable-5-1`.
- The Models tab refreshes automatically every five minutes while reachable. Its Refresh button, or `GET /v1/models?refresh=true`, requests fresh provider catalogs. See [model discovery](docs/api.md#get-v1models) for cache and availability limits.
- With multiple OpenAI accounts, `/v1/models` is the union of their catalogs. Balancing, conversation affinity, and failover stay within each model's supporting account subset. The Models tab shows account coverage and labels entries the upstream hides from its own picker.
- Subscription bars report upstream allowances, including Claude's separately scoped Fable weekly cap when present. Credit details and snapshot times appear under each account's **more** affordance; percentages are not token counts. See [Subscription Status](docs/subscription-status.md).
- AIRelays rejects requests when the selected runtime is disabled or the route is outside that runtime's published subset

## What AIRelays Exposes

- `GET /v1/models`
- `GET /v1/subscription/status` (OpenAI and `?provider=claude`)
- `GET /v1/account/rate_limits`
- `GET /v1/relay/status`
- `POST /v1/relay/accounts/refresh`
- `POST /v1/responses`
- `POST /v1/chat/completions`
- `POST /v1/completions`
- `POST /v1/files`
- `GET /v1/files`
- `GET /v1/files/{file_id}`
- `GET /v1/files/{file_id}/content`
- `DELETE /v1/files/{file_id}`
- `POST /v1/conversations`
- `GET /v1/conversations/{conversation_id}`
- `POST /v1/conversations/{conversation_id}`
- `DELETE /v1/conversations/{conversation_id}`
- `/no-tools/v1/models`
- `/no-tools/v1/responses`
- `/no-tools/v1/chat/completions`
- `/no-tools/v1/completions`

## What The Relay Changes (Compatibility Layer)

The ChatGPT subscription backend is not the public OpenAI platform API, so
AIRelays adapts some requests instead of letting them fail. Parameter
stripping is always visible: removed parameters are logged as a
`compatibility_adaptation` record in the traffic logs and reported in the
`x-airelays-ignored-parameters` response header. Other compatibility
normalizations are documented below and may log dedicated adaptation
records when the request shape itself is repaired.

**Sampling parameters are removed.** The upstream rejects `temperature`,
`top_p`, `presence_penalty`, and `frequency_penalty` outright
(`"Unsupported parameter: temperature"`). AIRelays strips them so standard
SDK calls keep working; generation then runs with the upstream's own
sampling defaults, which cannot be overridden. The Claude routes
apply the same adaptation — the local `claude` CLI has no sampling
controls — so the same SDK calls work against `claude:*` models too.

**Output-token limits are removed.** The subscription upstreams do not
honor client-set output caps (`max_tokens`, `max_completion_tokens`,
`max_output_tokens`), so AIRelays strips them instead of failing the
request; responses run to the model's natural stop.

**Cursor BYOK chat-route compatibility is normalized locally.** Some
current Cursor builds send malformed OpenAI-compatible requests to
`/v1/chat/completions`: either full Responses-style bodies (`input`,
`instructions`, flat `tools`, etc.) or flat Responses-style `custom`
tools/tool choices/tool calls such as `ApplyPatch`. AIRelays accepts those
shapes on the chat route, normalizes them to the Responses upstream, and
translates upstream `custom_tool_call` items back to chat-completions
`tool_calls` on both streaming and non-streaming responses. Requests that
mix `messages` and `input`, or tool outputs that do not reference a
preceding assistant tool call in the same request, are rejected loudly.

**Reasoning effort is forwarded, not invented.** `reasoning_effort` (chat
completions) and `reasoning: {"effort": ...}` (responses) pass through to
the upstream unchanged. Every model's supported modes are published in
`/v1/models` under `airelays.reasoning`, using provider catalog metadata
when available. Modes vary by model and may include `max` or `ultra`.
Claude modes map to the local CLI's `--effort` flag; unsupported values
are rejected with the model's supported list. When omitted, the provider
chooses its default; Claude can use an adaptive default. To choose
reasoning depth explicitly:

```bash
curl http://127.0.0.1:8317/v1/chat/completions \
  -H 'authorization: Bearer YOUR_AIRELAYS_TOKEN' \
  -H 'content-type: application/json' \
  -d '{
    "model": "gpt-5.5",
    "reasoning_effort": "medium",
    "messages": [{"role": "user", "content": "..."}]
  }'
```

**Conversations stick to one account.** With multiple OpenAI accounts, a
conversation keeps using the account that served its first turn (preserving
upstream prompt caching); it only fails over to another account at a turn
boundary when the pinned account is at its limit.

**Failed calls are retried automatically.** Transient upstream failures
(e.g. `server_is_overloaded`) are retried with exponential backoff — by
default 3 retries waiting 5s/20s/60s, each re-running account failover —
as long as no response byte has reached the client. A retry that succeeds
returns the normal response; a request that keeps failing returns
OpenAI-shaped error JSON with the real HTTP status and the upstream's own
reason. Tune with `retry_attempts` / `retry_backoff_seconds`
(`[providers.openai]`, desktop Settings → Providers; `0` disables).
Retries appear in the traffic log as `retry_backoff` records.

## Compatibility Boundary

OpenAI runtime:

- first-class routes: `/v1/responses`, `/v1/chat/completions`, `/v1/completions`
- local files and local conversations are supported
- non-stream responses are reconstructed from streamed upstream events
- `store=true` is rejected
- output-token limit fields are rejected explicitly on the OpenAI-shaped text-generation routes

Claude runtime:

- discovered Claude aliases and concrete model ids, plus configured overrides
- supported routes: text `/v1/chat/completions` and text `/v1/completions`
- stateless only
- no `/v1/responses`
- no files, images, audio, or tools
- structured outputs on chat completions: `response_format` `json_schema` /
  `json_object` map to the CLI's native `--json-schema` enforcement
- `reasoning_effort` maps to the CLI's `--effort` (low, medium, high, xhigh, max)
- no AIRelays local conversation reuse
- sampling parameters are stripped and disclosed, like on the OpenAI runtime

## Security Defaults

- default listener: `127.0.0.1:8080`
- protected routes: `/v1/*` and `/no-tools/v1/*`
- public routes: `/` and `GET /healthz`
- protected diagnostics: `GET /v1/relay/status`
- default rate limit: `120` requests/minute with burst `40`
- default concurrent request cap: `8` per IP
- repeated bad tokens trigger a temporary IP block
- the Claude runtime is loopback-only and follows the relay's protected or open local auth mode

## Configuration

Traffic logs are bounded by default: up to **7 days**, **1 GiB total**, and
**50 MiB per file**, with hourly rotation and oldest-first cleanup. A busy
relay may keep less history because the disk limit takes precedence.
Use `airelays logs --retention-days 30` for a month, or the tray app's
**Settings → Traffic log retention** controls. Changes persist without a
restart; `GET`/`PUT /v1/relay/logging` expose the same policy and current usage.
Starting the upgraded relay also applies the limits to existing traffic logs.
See [retention configuration](docs/configuration.md#traffic-log-retention).

AIRelays reads configuration in this order:

1. CLI flags
2. `AIRELAYS_*` environment variables
3. legacy `OPENAI_ENDPOINT_*` migration variables where supported
4. `~/.config/airelays/config.toml`
5. built-in defaults

Important toggles:

- `AIRELAYS_REQUIRE_BEARER_AUTH`
- `AIRELAYS_BEARER_TOKEN`
- `AIRELAYS_BEARER_TOKEN_FILE`
- `AIRELAYS_ENABLE_OPENAI`
- `AIRELAYS_OPENAI_MODELS_CACHE_TTL_SECONDS`
- `AIRELAYS_ENABLE_CLAUDE`
- `AIRELAYS_CLAUDE_BIN`
- `AIRELAYS_CLAUDE_MODELS`

## Paths

- config: `~/.config/airelays/config.toml`
- data dir: `~/.airelays`
- logs: `~/.airelays/logs`
- relay token: `~/.airelays/relay-token`
- earlier singular AIRelay paths remain compatible for local upgrades

## More Docs

- [docs/getting-started.md](docs/getting-started.md)
- [docs/configuration.md](docs/configuration.md)
- [docs/security.md](docs/security.md)
- [docs/api.md](docs/api.md)
- [docs/architecture.md](docs/architecture.md)
- [docs/subscription-status.md](docs/subscription-status.md)
- [docs/faq.md](docs/faq.md)
- [docs/troubleshooting.md](docs/troubleshooting.md)
- [docs/disclaimer.md](docs/disclaimer.md)
