Metadata-Version: 2.5
Name: spatialreal
Version: 0.1.0
Summary: SpatialReal Python SDK — realtime avatar sessions (audio in, animation/egress out)
Project-URL: Documentation, https://docs.spatialreal.com
Project-URL: Website, https://www.spatialreal.ai/
Project-URL: Source, https://github.com/SpatialReal-ai/python-sdk
Author-email: SpatialReal <hello@spatialreal.com>
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: audio,avatar,realtime,spatialreal,websocket
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Topic :: Multimedia :: Sound/Audio
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10.0
Requires-Dist: aiohttp>=3.9
Requires-Dist: nanoid>=2.0
Requires-Dist: protobuf<8,>=6.31
Requires-Dist: websockets>=12
Provides-Extra: dev
Requires-Dist: grpcio-tools==1.75.1; extra == 'dev'
Requires-Dist: pytest; extra == 'dev'
Requires-Dist: pytest-asyncio; extra == 'dev'
Requires-Dist: ruff; extra == 'dev'
Description-Content-Type: text/markdown

# SpatialReal Python SDK

`pip install spatialreal` — realtime avatar sessions against the SpatialReal
backend: stream audio in, get lifecycle signals back, and (in egress mode) have
the avatar's synchronized audio/video published straight into a LiveKit room or
Agora channel.

## Usage

```python
from datetime import datetime, timedelta, timezone
from spatialreal import LiveKitEgressConfig, new_avatar_session

session = new_avatar_session(
    api_key="...",
    app_id="...",
    avatar_id="...",
    console_endpoint_url="https://api.spatialreal.cloud",  # OpenAPI root: session tokens
    ingress_endpoint_url="wss://driven.us-west.spatialreal.cloud/v2/driveningress",
    expire_at=datetime.now(timezone.utc) + timedelta(hours=1),
    sample_rate=16000,
    livekit_egress=LiveKitEgressConfig(url=..., api_token=..., room_name=..., publisher_id=...),
    on_playback=lambda sig: print(sig.req_id, "ended" if sig.end else "playing"),
    on_error=lambda err: print("error:", err, "retryable:", getattr(err, "retryable", None)),
    on_close=lambda: print("closed"),
)
await session.init()  # API key -> session token (console API)
await session.start()  # WebSocket + handshake; read loop starts

req_id = await session.send_audio(pcm_bytes)  # same id until end=True
await session.send_audio(b"", end=True)  # closes the segment
await session.interrupt()  # interrupts the latest request
await session.close()
```

### Pause / resume (egress mode)

When the server declares the `playback_control` capability, the current
segment can be paused and resumed instead of interrupted — the server holds
its send cursor with zero regeneration:

```python
await session.pause()  # server stops writing frames, keeps ingesting audio
# ...keep sending audio, including end=True, while paused...
await session.resume()  # continues from where it stopped
```

Feature-detect with `"playback_control" in session.capabilities`. Subscribe to
`on_playback_state(PlaybackStateEvent)` for the resulting state (PLAYING /
PAUSED / ENDED / INTERRUPTED with `played_ms` and, on interrupt, a `reason` such
as `pause_timeout`). Older servers never send it and ignore pause/resume.

Session objects are single-use: on a dropped connection, create a fresh one
(request state is deliberately not reusable across connections).
`AvatarSDKError.retryable` tells you whether reconnecting can succeed — it
implements the server's WebSocket close-code contract (40xx: don't retry,
45xx: retry with backoff).

## What's new

- `on_playback(PlaybackSignal)` — structured playback lifecycle; no protobuf
  parsing in caller code (`transport_frames(raw, is_last)` still exists).
- `CloseCode` + `AvatarSDKError.retryable` — the reconnect contract as API.
- `session.capabilities` — server-declared capabilities from the handshake
  (feature-detect, don't version-detect).
- `pause()` / `resume()` + `on_playback_state(PlaybackStateEvent)` — egress-mode
  server-side playback control (gated on the `playback_control` capability).
- `on_close` fires exactly once per session; the token request has a timeout.
- `AudioFormat.OGG_OPUS` takes pre-encoded bytes; the SDK does not encode
  client-side.

## Development

```bash
uv venv && uv pip install -e ".[dev]"
python -m pytest tests -q     # protocol-level tests against an in-process fake backend
./scripts/gen_proto.sh        # regenerate protobuf from proto/ (see proto/SHARED_PROTO_COMMIT)
```

## License

Apache-2.0
