Metadata-Version: 2.5
Name: mlx-beam
Version: 0.1.0a2
Summary: Light and modular inference engine, built on MLX.
Project-URL: Homepage, https://p4ik.github.io/mlx-beam/
Project-URL: Repository, https://github.com/p4ik/mlx-beam
Project-URL: Changelog, https://github.com/p4ik/mlx-beam/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/p4ik/mlx-beam/issues
Author: the mlx-beam contributors
License-Expression: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: apple-silicon,inference,llm,mlx,moe
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: MacOS
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.11
Requires-Dist: huggingface-hub
Requires-Dist: jinja2
Requires-Dist: mlx>=0.32; sys_platform == 'darwin'
Requires-Dist: numpy
Requires-Dist: protobuf
Requires-Dist: regex
Requires-Dist: sentencepiece
Requires-Dist: transformers>=5.7
Provides-Extra: dev
Requires-Dist: mlx[cpu]>=0.32; (sys_platform == 'linux') and extra == 'dev'
Requires-Dist: pre-commit>=4; extra == 'dev'
Requires-Dist: pytest>=8; extra == 'dev'
Description-Content-Type: text/markdown

# mlx-beam

**B.E.A.M. — Batched Engine for Apple Metal.** Light and modular inference engine, built on [MLX](https://github.com/ml-explore/mlx).

> **Work in progress.** See the status table below and the [changelog](https://github.com/p4ik/mlx-beam/blob/main/CHANGELOG.md).

## Install

```bash
uv tool install mlx-beam
beam doctor
```

Inside a uv project: `uv add mlx-beam`, then `uv run beam doctor`.

`beam doctor` reports the Python, MLX, device and memory it sees (`--json` for scripts) and exits non-zero when MLX is missing or fails to load.

## Serve

```bash
beam serve --model p4ik/Qwen3.8-27B-MLX-OptiQ-5bit --port 8000 \
  --max-completion-tokens 4096 --min-response-tokens 512 --max-reasoning-tokens 8192 \
  --kv-bits 8
```

That is an OpenAI-compatible server (`/v1/chat/completions`, `/v1/completions`, `/v1/responses`, `/v1/models`) plus `/health`, which reports what was actually built: the KV layout per layer, the batching and cache settings, and every request default with where it came from (flag, the model's `generation_config.json`, or mlx-lm's own).

Flags follow mlx-lm's names where mlx-lm has one (`--temp`, `--top-p`, `--kv-bits`, `--prompt-cache-size`, `--chat-template`, …). Token limits say what they count: `--max-context` (prompt plus generated, a hard cap), `--max-prompt-tokens` (prompt, a hard cap), `--max-completion-tokens` (generated, the default a request may override), `--max-reasoning-tokens` (the think block; closed by force at the budget) and `--min-response-tokens` (what the answer keeps after the block). `beam serve --help` lists them all with their units.

Requests may use the names other servers taught clients: `max_tokens`, `thinking_token_budget`, `reasoning: {effort, max_tokens}`, `enable_thinking`, `reasoning_effort`. The model's thinking is returned in `reasoning` (`--reasoning-field` switches to `reasoning_content`, both, or none), counted in `usage.completion_tokens_details.reasoning_tokens`, and flagged there when a limit cut it (`thinking_truncated`, `response_truncated`).

## What sets it apart

- **Robust prefix cache** — RAM and SSD tiers, checkpoints for hybrid models. Survives model swaps and restarts.
- **Expert streaming** — Mixture-of-experts models larger than memory. Residency configurable, from minimal RAM to fully resident.
- **No bloat** — The core is the token path. Vision, audio, conversion, structured output and tool-call repair are optional extras.
- **Batched MTP** — Multi-token prediction stays on with many requests at once.
- **Batched vision** — Images go through the same scheduler; no request waits behind a picture.
- **No stalls** — A short request beside a long prefill answers in seconds.
- **Mixed-precision KV cache** — Bits per layer, set at conversion.
- **Thinking budget** — A hard cap on the reasoning trace, per request.
- **Responses API** — Next to chat completions, stateless.

The engine reads standard MLX checkpoints and the B.E.A.M. package layout (`extras/` next to the shards; see the model cards under [huggingface.co/p4ik](https://huggingface.co/p4ik)).

## Why it exists

Existing MLX servers either stop at the basics or grow things that have no place in an inference engine: a built-in game, a cloud path that arrives with an update. The ones we ran daily also had bugs where it matters most: prefix cache, batching under load, vision. B.E.A.M. keeps the core to the token path and fixes those paths at the source. Everything else is an extra you choose to install; nothing ever ships in the core that you did not ask for.

## Status

| Piece | State |
|---|---|
| CLI, packaging, CI | done |
| Vendored mlx-lm base (pinned, four local changes) | done |
| OpenAI-compatible server, continuous batching, quantized KV cache | done, text only |
| Prefix cache with recurrent-state checkpoints | done, RAM tier |
| Reasoning budget, request defaults, sampling controls | done |
| Multi-token prediction in the batch | planned |
| Vision, structured output | planned, as extras |
| Expert streaming from SSD | planned |

Measured numbers are published as they are measured, with machine, model and date.

## Contributing

See [CONTRIBUTING.md](https://github.com/p4ik/mlx-beam/blob/main/CONTRIBUTING.md). Rules for coding agents are in [AGENTS.md](https://github.com/p4ik/mlx-beam/blob/main/AGENTS.md).

## License

Apache-2.0. Vendored components keep their own licenses; see [NOTICE](https://github.com/p4ik/mlx-beam/blob/main/NOTICE).
