Metadata-Version: 2.5
Name: hfl
Version: 0.27.0
Summary: Download, run and try any Hugging Face model on your own machine.
Project-URL: Homepage, https://ggalancs.github.io/hfl/
Project-URL: Documentation, https://ggalancs.github.io/hfl/
Project-URL: Repository, https://github.com/ggalancs/hfl
Project-URL: Issues, https://github.com/ggalancs/hfl/issues
Project-URL: Changelog, https://github.com/ggalancs/hfl/blob/main/CHANGELOG.md
Author: Gabriel Galán Pelayo
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: ai,gguf,huggingface,inference,llama-cpp,llm,local-llm,mlx,model-serving,ollama,openai-api,self-hosted,transformers,vllm
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Framework :: FastAPI
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Requires-Dist: fastapi<0.143,>=0.135.0
Requires-Dist: httpx~=0.28.0
Requires-Dist: huggingface-hub<2.0,>=1.5.0
Requires-Dist: psutil<8,>=5.9
Requires-Dist: pydantic<2.14,>=2.12.0
Requires-Dist: python-multipart>=0.0.31
Requires-Dist: pyyaml~=6.0
Requires-Dist: rich<16,>=13.0
Requires-Dist: starlette<2.0,>=1.3.1
Requires-Dist: typer<1.0,>=0.15.0
Requires-Dist: uvicorn[standard]<0.55,>=0.34.0
Provides-Extra: all
Requires-Dist: accelerate<2.0,>=1.13.0; extra == 'all'
Requires-Dist: accelerate<2.0,>=1.2.0; extra == 'all'
Requires-Dist: av<19,>=11; extra == 'all'
Requires-Dist: bitsandbytes<1.0,>=0.45.0; (sys_platform != 'darwin') and extra == 'all'
Requires-Dist: coqui-tts[codec]<1.0,>=0.24.0; extra == 'all'
Requires-Dist: diffusers<1.0,>=0.30.0; extra == 'all'
Requires-Dist: faster-whisper<2.0,>=1.0.0; extra == 'all'
Requires-Dist: gguf>=0.10.0; extra == 'all'
Requires-Dist: llama-cpp-python>=0.3.20; (sys_platform != 'win32') and extra == 'all'
Requires-Dist: llguidance<2.0,>=1.7.0; extra == 'all'
Requires-Dist: llguidance<2.0,>=1.7.0; (sys_platform == 'darwin' and platform_machine == 'arm64') and extra == 'all'
Requires-Dist: mcp<3.0,>=1.0.0; extra == 'all'
Requires-Dist: mlx-lm<1.0,>=0.31.2; (sys_platform == 'darwin' and platform_machine == 'arm64') and extra == 'all'
Requires-Dist: peft<1.0,>=0.17.0; extra == 'all'
Requires-Dist: pillow>=12.3.0; extra == 'all'
Requires-Dist: pystray>=0.19.0; extra == 'all'
Requires-Dist: safetensors<1.0,>=0.4.0; extra == 'all'
Requires-Dist: sentencepiece>=0.2.0; extra == 'all'
Requires-Dist: sounddevice>=0.4.6; extra == 'all'
Requires-Dist: soundfile>=0.12.1; extra == 'all'
Requires-Dist: torch<3.0,>=2.5.0; extra == 'all'
Requires-Dist: torchaudio<3.0,>=2.2.0; extra == 'all'
Requires-Dist: transformers<6.0,>=5.0.0; extra == 'all'
Requires-Dist: vllm<1.0,>=0.30.0; (sys_platform == 'linux') and extra == 'all'
Provides-Extra: audio
Requires-Dist: sounddevice>=0.4.6; extra == 'audio'
Requires-Dist: soundfile>=0.12.1; extra == 'audio'
Provides-Extra: build
Requires-Dist: pyinstaller>=6.0; extra == 'build'
Provides-Extra: convert
Requires-Dist: gguf>=0.10.0; extra == 'convert'
Provides-Extra: coqui
Requires-Dist: coqui-tts[codec]<1.0,>=0.24.0; extra == 'coqui'
Requires-Dist: soundfile>=0.12.1; extra == 'coqui'
Requires-Dist: torch<3.0,>=2.5.0; extra == 'coqui'
Requires-Dist: torchaudio<3.0,>=2.2.0; extra == 'coqui'
Provides-Extra: dev
Requires-Dist: mypy>=1.10.0; extra == 'dev'
Requires-Dist: numpy>=1.24.0; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.24.0; extra == 'dev'
Requires-Dist: pytest-cov>=4.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.8.0; extra == 'dev'
Provides-Extra: imagegen
Requires-Dist: accelerate<2.0,>=1.2.0; extra == 'imagegen'
Requires-Dist: diffusers<1.0,>=0.30.0; extra == 'imagegen'
Requires-Dist: pillow>=12.3.0; extra == 'imagegen'
Requires-Dist: safetensors<1.0,>=0.4.0; extra == 'imagegen'
Requires-Dist: torch<3.0,>=2.5.0; extra == 'imagegen'
Provides-Extra: llama
Requires-Dist: llama-cpp-python>=0.3.20; extra == 'llama'
Provides-Extra: mcp
Requires-Dist: mcp<3.0,>=1.0.0; extra == 'mcp'
Provides-Extra: mlx
Requires-Dist: llguidance<2.0,>=1.7.0; (sys_platform == 'darwin' and platform_machine == 'arm64') and extra == 'mlx'
Requires-Dist: mlx-lm<1.0,>=0.31.2; (sys_platform == 'darwin' and platform_machine == 'arm64') and extra == 'mlx'
Provides-Extra: otel
Requires-Dist: opentelemetry-api>=1.27.0; extra == 'otel'
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.27.0; extra == 'otel'
Requires-Dist: opentelemetry-sdk>=1.27.0; extra == 'otel'
Provides-Extra: rocm
Requires-Dist: llama-cpp-python>=0.3.20; extra == 'rocm'
Provides-Extra: structured
Requires-Dist: llguidance<2.0,>=1.7.0; extra == 'structured'
Provides-Extra: stt
Requires-Dist: av<19,>=11; extra == 'stt'
Requires-Dist: faster-whisper<2.0,>=1.0.0; extra == 'stt'
Provides-Extra: train
Requires-Dist: accelerate<2.0,>=1.13.0; extra == 'train'
Requires-Dist: bitsandbytes<1.0,>=0.45.0; (sys_platform != 'darwin') and extra == 'train'
Requires-Dist: llguidance<2.0,>=1.7.0; extra == 'train'
Requires-Dist: peft<1.0,>=0.17.0; extra == 'train'
Requires-Dist: sentencepiece>=0.2.0; extra == 'train'
Requires-Dist: torch<3.0,>=2.5.0; extra == 'train'
Requires-Dist: transformers<6.0,>=5.0.0; extra == 'train'
Provides-Extra: transformers
Requires-Dist: accelerate<2.0,>=1.13.0; extra == 'transformers'
Requires-Dist: bitsandbytes<1.0,>=0.45.0; (sys_platform != 'darwin') and extra == 'transformers'
Requires-Dist: llguidance<2.0,>=1.7.0; extra == 'transformers'
Requires-Dist: sentencepiece>=0.2.0; extra == 'transformers'
Requires-Dist: torch<3.0,>=2.5.0; extra == 'transformers'
Requires-Dist: transformers<6.0,>=5.0.0; extra == 'transformers'
Provides-Extra: tray
Requires-Dist: pillow>=12.3.0; extra == 'tray'
Requires-Dist: pystray>=0.19.0; extra == 'tray'
Provides-Extra: tts
Requires-Dist: soundfile>=0.12.1; extra == 'tts'
Requires-Dist: torch<3.0,>=2.5.0; extra == 'tts'
Requires-Dist: torchaudio<3.0,>=2.2.0; extra == 'tts'
Requires-Dist: transformers<6.0,>=5.0.0; extra == 'tts'
Provides-Extra: vllm
Requires-Dist: vllm<1.0,>=0.30.0; (sys_platform == 'linux') and extra == 'vllm'
Provides-Extra: vulkan
Requires-Dist: llama-cpp-python>=0.3.20; extra == 'vulkan'
Description-Content-Type: text/markdown

<div align="center">

# HFL

**Download, run and try any Hugging Face model on your own machine.**

One command takes a model from the Hub to a local chat or to an OpenAI-,
Ollama- and Anthropic-compatible API. No account. No cloud of its own — by design.

[![PyPI](https://img.shields.io/pypi/v/hfl.svg)](https://pypi.org/project/hfl/)
[![Python](https://img.shields.io/pypi/pyversions/hfl.svg)](https://pypi.org/project/hfl/)
[![License: Apache 2.0](https://img.shields.io/badge/License-Apache--2.0-blue.svg)](https://github.com/ggalancs/hfl/blob/main/LICENSE)
[![Docker](https://img.shields.io/badge/docker-ghcr.io%2Fggalancs%2Fhfl-2496ED.svg)](https://github.com/ggalancs/hfl/pkgs/container/hfl)
[![CI](https://github.com/ggalancs/hfl/actions/workflows/ci.yml/badge.svg)](https://github.com/ggalancs/hfl/actions/workflows/ci.yml)

[Quick start](#quick-start) · [Why HFL](#why-hfl) · [Install](#install) · [Use it](#use-it) · [API](#connect-your-tools) · [Docs](#documentation) · **[Español](https://github.com/ggalancs/hfl/blob/main/README.es.md)**

</div>

<p align="center">
  <img src="https://raw.githubusercontent.com/ggalancs/hfl/main/docs/assets/hfl-run-demo.svg" alt="Terminal, animated: hfl run downloads a model from the Hugging Face Hub on first use, checks it against the Hub's sha256, loads it and answers a question (real output, abridged, waits shortened)" width="880">
</p>

## Quick start

```bash
pip install "hfl[llama,mlx]"        # the MLX part installs only on Apple Silicon

hfl start                            # suggests a model that fits this machine, then chats
hfl run hf.co/bartowski/Llama-3.2-3B-Instruct-GGUF:Q4_K_M   # or any model on the Hub
```

`hfl start` looks at this machine, offers up to three small Apache-2.0 models
that fit (no sign-up, no terms to accept), downloads the one you pick and opens
a chat: from nothing to a first answer in about 30 seconds on a Mac. Any model is
downloaded the first time and reused from disk after that. To serve it instead:

```bash
hfl serve --model hf.co/bartowski/Llama-3.2-3B-Instruct-GGUF:Q4_K_M
```

Any OpenAI, Ollama or Anthropic client can now talk to `http://localhost:11434` — and
opening that address in a browser gives you a chat page, served by HFL itself
(no account, nothing loaded from the internet): your conversations, a system
prompt and temperature per conversation, and images for models that can see
them. Its model panel searches the Hub, downloads (with progress) and deletes
models — owner operations, so the page is allowed them only once you start HFL
with `HFL_ORIGINS=http://127.0.0.1:11434` (that trusts every page served at
that address, including `/docs`, which loads scripts from a CDN).

## Why HFL

- **The whole Hub, not one file format.** GGUF repos run through llama.cpp, MLX
  builds run natively on Apple Silicon, and safetensors checkpoints are converted
  and quantized for you on pull. Copy a repo name from the Hub and run it.
- **Browse the Hub from your terminal.** `hfl search` pages through the live
  Hub with sizes, downloads and formats at a glance; ask in plain words
  ("coding assistant 7b") or filter by GGUF or by size,
  press a number and the model is pulled, license check included.
- **Yours alone.** No account, no sign-in, no cloud service behind it. On its own
  HFL talks to one server, the Hugging Face Hub, to fetch weights; its web-search
  endpoints reach the web only when a client calls them, and a test pins every
  host the code can reach. With no network, everything you have already pulled
  keeps working.
- **Plugs into what you already use.** OpenAI (chat, completions, embeddings,
  Responses), Ollama and Anthropic Messages APIs on one port, with structured
  tool calling for Qwen, Llama 3, Mistral, Gemma 4, gpt-oss, DeepSeek, GLM and
  Hermes families and JSON-schema outputs.
  Existing `OLLAMA_*` settings such as `OLLAMA_HOST` are honoured.
- **As many models as your memory holds.** Before every load HFL estimates what
  the model will take (weights + KV cache) and keeps the machine under a memory
  budget you set: models load side by side while they fit, idle ones make room,
  one in use is never pulled out from under a request, and one that cannot fit
  is refused with the numbers — before anything is unloaded. GPU-aware on NVIDIA and AMD.
- **Knows the Hub.** Find models that fit your hardware (`hfl recommend`), pick
  the best community quant for your machine (`hfl pull-smart`), size
  Mixture-of-Experts models by their total parameters, check licenses before
  downloading and keep a provenance record of every pull.

### Compared with downloading the files yourself

`hf download` and `git lfs clone` fetch a repo's files; what to do with them
is left to you. HFL fetches only the files a model needs and runs them:

| | `hf download` / `git lfs` | HFL |
|---|---|---|
| Which files | the whole repo, or the patterns you give | the one quant you asked for (plus a split model's parts and a vision projector) |
| safetensors checkpoints | downloaded as they are | converted to GGUF and quantized, or run as they are (MLX, Transformers) |
| Checked after download | `hf download`: the size; `git lfs`: each object's sha256 | against the Hub's sha256; a damaged file is fetched again |
| License | not looked at | checked before downloading, and kept with the model with the repo and commit it came from |
| Then | your own code or server | a chat, and OpenAI, Ollama and Anthropic APIs on one port |

## Install

| How | Command |
|---|---|
| **pip** (recommended) | `pip install "hfl[llama,mlx]"`, then `hfl install llama-server` (llama.cpp's official build: GGUF models answer 4 requests at once) |
| **pip on Windows** | `pip install hfl` then `hfl install llama-server` (no compiler needed; `hfl[llama]` needs Visual Studio's C++ Build Tools) |
| **Docker** | `docker run -p 127.0.0.1:11434:11434 -v hfl:/var/lib/hfl ghcr.io/ggalancs/hfl` (this machine only; to open it to the network, publish `-p 11434:11434` with `-e HFL_API_KEY=…`). llama-server included: GGUF models answer 4 requests at once |
| **Installers** | `.dmg`, `.msi` and standalone binaries on the [releases page](https://github.com/ggalancs/hfl/releases), llama-server included |
| **From source** | `git clone https://github.com/ggalancs/hfl && cd hfl && pip install -e ".[llama,mlx]"` |

<details>
<summary><b>Optional extras</b> — GPU, speech, vLLM and more</summary>

| Extra | Adds |
|---|---|
| `llama` | llama.cpp for GGUF models (Metal on Apple Silicon out of the box) |
| `mlx` | Native MLX backend on Apple Silicon |
| `transformers` | Transformers backend for GPU inference with bitsandbytes |
| `vllm` | vLLM backend |
| `convert` | Tools to convert safetensors to GGUF |
| `tts` / `coqui` | Text-to-speech (Bark, SpeechT5 / Coqui XTTS, VITS) |
| `stt` | Speech-to-text (Whisper) |
| `mcp` | Model Context Protocol client and server |
| `all` | Everything above (on Windows without `llama`: see the install table) |

Converting safetensors to GGUF fetches llama.cpp's converter (Python) the
first time and quantizes with a `llama-quantize` already on the machine:
Homebrew's (`brew install llama.cpp`), a distro package's, or the one inside
llama-cpp-python (`pip install 'hfl[llama]'`). Nothing is compiled unless none
of those exists; then it needs **git**, **cmake** and a **C++ compiler**.
Pre-quantized GGUF and MLX models need none of this.

On an AMD GPU (ROCm, Linux x64), `hfl install llama-server --variant rocm`
installs llama.cpp's ROCm build for GGUF models. The `transformers` and
`vllm` extras install PyTorch's NVIDIA build by default, which cannot use an
AMD GPU: install PyTorch's ROCm build first, as
[pytorch.org](https://pytorch.org/get-started/locally/) shows for your ROCm
version.

</details>

## Use it

### Search and pick from the terminal

<p align="center">
  <img src="https://raw.githubusercontent.com/ggalancs/hfl/main/docs/assets/hfl-search-demo.svg" alt="Terminal: hfl search lists Hugging Face Hub models page by page with size, downloads and format; pressing a number pulls the model (real output)" width="880">
</p>

```bash
hfl search qwen3                          # everything matching, most downloaded first
hfl search qwen3 --gguf --max-params 8    # GGUF only, 8B or smaller — searched across the whole Hub
hfl search llama --sort likes             # or: downloads (default), created
hfl search "coding assistant 7b"         # read as: coding models of about 7B
```

Each page lists up to ten models with their size, downloads, likes, format and
task. Press **0–9** to pull one (you confirm, and its license is checked
first), **SPACE** for the next page, **p** for the previous one, **q** to leave.
Sizes are total parameters, so Mixture-of-Experts models are not shown smaller
than they are.

### Chat

```bash
hfl run hf.co/bartowski/Llama-3.2-3B-Instruct-GGUF:Q4_K_M   # from the Hub, pulled on first use
hfl run qwen3-coder                                          # a short name, looked up on the Hub
hfl run hf.co/mlx-community/Qwen2.5-0.5B-Instruct-4bit       # an MLX build on Apple Silicon
hfl run llama70b --system "You are a Python expert"          # a local name or alias
hfl run llama70b --session work                              # resume and save a conversation
```

The `hf.co/` prefix is optional. `:Q4_K_M` picks a quantization, `@<ref>` pins a
branch, tag or commit.

A short name (`qwen3-coder`, `llama3.2`, `qwen3:8b`, `gemma3:4b-q8_0`) that is
not a local model is looked up among the Hub's GGUF builds: instruct builds and
the usual quantizers first, derivatives (abliterated, merges...) left out, and a
quantization that fits your machine (Q4_K_M unless it does not). You pick from the
list before anything is downloaded (`--yes` takes the first), and the name is kept
as an alias, so the next `hfl run qwen3-coder` is local. Ollama clients get the
same over `/api/pull {"model": "llama3.2"}`.

`--session <name>` keeps the conversation: the next `hfl run <model> --session
<name>` picks it up where it stopped. Sessions are JSON files under
`~/.hfl/sessions/`:

```bash
hfl sessions list          # name, model, messages, last update
hfl sessions show work     # print a session's messages
hfl sessions rm work       # delete it
```

### Pull, search and manage

```bash
hfl pull meta-llama/Llama-3.3-70B-Instruct                 # Q4_K_M by default
hfl pull meta-llama/Llama-3.3-70B-Instruct --quantize Q5_K_M --alias llama70b
hfl pull meta-llama/Llama-3.3-70B-Instruct@a1b2c3d          # reproducible: pinned revision

hfl list                                # what is on this machine
hfl inspect llama70b                    # details and license
hfl outdated                            # newer versions on the Hub? (downloads nothing)
hfl rm llama70b

hfl import ~/.lmstudio/models/lmstudio-community/Qwen3-8B-GGUF   # a GGUF you already have
hfl import ~/.lmstudio/models/mlx-community/Qwen3-8B-4bit        # or MLX / Hugging Face weights
```

`hfl import` registers a model where it is — no copy, no server — for models
downloaded by LM Studio, llama.cpp, `huggingface-cli` or by hand: a GGUF (a
split model or a vision model's projector beside it are handled) or a folder of
MLX or Hugging Face weights (`config.json` and `.safetensors`). `hfl rm` never
deletes anything outside HFL's own folder.

### Find the right model

```bash
hfl recommend                           # top models that fit THIS machine's RAM/VRAM
hfl discover --family qwen              # filter the live Hub; marks what you already have
hfl pull-smart Qwen/Qwen3-30B-A3B       # best community variant for your hardware
hfl verify <model>                      # tokenizer, chat template, smoke generation, tools
hfl bench <model>                       # time to first token, tokens/s, p50/p95
```

See [docs/hub-native-features.md](https://github.com/ggalancs/hfl/blob/main/docs/hub-native-features.md) for every option.

### Several models at once

HFL keeps every model that fits under `HFL_MEMORY_BUDGET` — the share of total
RAM the machine may have in use after a load, other programs included (default
`85%`). Each load reports the numbers before it happens:

```text
Memory: 65.3 of 128.0 GB in use (51%). qwen3-14b needs ~9.0 GB → after loading, 74.3 GB in use (58%); budget 85%.
```

- fits → it loads next to the models already resident;
- does not fit → idle models are unloaded, least recently used first;
- the room is held by models answering requests → the load waits for them;
- cannot fit even alone → refused with the numbers (HTTP 507), nothing unloaded.

With an NVIDIA or AMD GPU the model must also fit the card (read through `nvidia-smi`,
or `rocm-smi`/`amd-smi` on ROCm).
Idle models unload after `keep_alive` (default `5m`, renewed on every use).
`hfl ps` and `GET /api/ps` show what is loaded and how much room is left.

<details>
<summary><b>Tool calling</b> — agents work out of the box</summary>

Send `tools` on `/api/chat`, `/v1/chat/completions` or `/v1/messages`: HFL
renders them through the model's own chat template (Qwen and Hermes
`<tool_call>`, Llama 3 `<|python_tag|>`, Mistral `[TOOL_CALLS]`, gpt-oss's
Harmony channels, DeepSeek's and GLM's own markers), parses the reply into
`message.tool_calls` with the arguments as an object, and accepts
`role: "tool"` results on the next turn. When a model's template has no place
for tools (Hermes-3, DeepSeek-R1), HFL writes them into the system prompt in
the Hermes convention.

```bash
curl http://localhost:11434/api/chat -d '{
  "model": "qwen3-32b-q4_k_m",
  "stream": false,
  "messages": [{"role": "user", "content": "Save Hello at topics/hello.md"}],
  "tools": [{"type": "function", "function": {
    "name": "write_wiki", "description": "Create or overwrite a wiki article",
    "parameters": {"type": "object",
      "properties": {"path": {"type": "string"}, "content": {"type": "string"}},
      "required": ["path", "content"]}}}]
}'
```

```json
{"message": {"role": "assistant", "content": "",
  "tool_calls": [{"function": {"name": "write_wiki",
    "arguments": {"path": "topics/hello.md", "content": "Hello"}}}]},
 "done": true}
```

When streaming, `tool_calls` arrive on the final `done: true` chunk. The
executable spec lives in `tests/test_tool_calling_acceptance.py`.

</details>

<details>
<summary><b>Text to speech</b> — Bark, SpeechT5, Coqui XTTS</summary>

```bash
pip install "hfl[tts,audio]"
hfl pull suno/bark-small --alias bark
hfl tts bark "Hello, this is a test." -o hello.wav     # write a file (wav, mp3, ogg)
hfl speak bark "Hola mundo" --lang es --speed 0.9      # play it
```

Options: `--lang`, `--voice`, `--speed` (0.25–4.0), and for `tts` also
`--output`, `--rate` and `--format`. The same voices are served over HTTP:

```bash
# OpenAI-compatible
curl http://localhost:11434/v1/audio/speech -H "Content-Type: application/json" \
  -d '{"model": "bark", "input": "Hello world", "voice": "alloy"}' --output speech.wav

# Native: language, speed, sample rate and format (wav, mp3, ogg)
curl http://localhost:11434/api/tts -H "Content-Type: application/json" \
  -d '{"model": "bark", "text": "Hola mundo", "language": "es"}' --output speech.wav
```

</details>

<details>
<summary><b>More tools</b> — LoRA, KV snapshots, speculative decoding, MCP, Hub upload</summary>

- `hfl lora apply|remove|list` — hot-swap LoRA adapters without reloading the base model. With `--parallel` (llama-server), its process starts again with the new set, once no reply is in progress.
- `hfl snapshot save|load|list|delete` — persist the KV cache to disk for warm starts.
- `hfl draft-recommend` — pick a small Hub sibling for speculative decoding.
- `hfl train <model> --data data.jsonl` — train a LoRA adapter on a local safetensors model
  (MLX on Apple Silicon, `[mlx]`; Transformers + PEFT elsewhere, `[train]` — there the
  adapter comes merged); the result is a model of its own. `--fuse` merges it, `--gguf Q4_K_M`
  also exports it for llama.cpp. Also `POST /api/train` (owner, on the host).
- `hfl mcp serve` / `hfl mcp connect` — run as a Model Context Protocol server, or use MCP tools.
- `hfl compliance-dashboard` — license risk across your local models.
- `POST /api/push` — upload a registered model to the Hub.
- `WS /ws/chat` — bidirectional chat with frame-level cancellation.

</details>

## Connect your tools

**Coding agents.** One command opens Claude Code or Codex on a local model:

```bash
hfl launch claude -m hf.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF:Q4_K_M
hfl launch codex  -m qwen-coder                 # a local name or alias works too
hfl launch claude -m qwen-coder --print         # just show the settings for your shell
```

HFL downloads the model if needed, starts a server if none is running (and
stops it when the agent exits), loads the model and hands the agent its real
context window. Nothing in the agent's own configuration is changed.
Arguments after `--` go to the agent: `hfl launch claude -m qwen-coder -- -p "fix the tests"`.

The server listens on `http://localhost:11434` and speaks three APIs. With
llama.cpp's `llama-server` (`hfl install llama-server` fetches its official
build; a `brew install llama.cpp` serves too), it serves each
GGUF model with 4 parallel slots by default — what coding agents and several
users need; two models answer at the same time too. Without it, or with
`HFL_NUM_PARALLEL=1`, each GGUF model answers one request at a time in process.
On Apple Silicon, MLX models serve 4 requests at once too (mlx-lm's batching,
nothing to install): measured with Qwen3-14B, 4x the throughput with 8 clients.

**OpenAI** — `/v1/chat/completions`, `/v1/completions`, `/v1/embeddings`, `/v1/responses`,
`/v1/audio/speech`, `/v1/audio/transcriptions`, `/v1/images/generations`

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="not-needed")
reply = client.chat.completions.create(
    model="hf.co/bartowski/Llama-3.2-3B-Instruct-GGUF:Q4_K_M",
    messages=[{"role": "user", "content": "Explain quantum computing in one paragraph"}],
)
print(reply.choices[0].message.content)
```

**Ollama** — `/api/chat`, `/api/generate`, `/api/embed`, `/api/tags`, `/api/ps`, `/api/pull`,
`/api/delete` (only from the server's own machine) and more

```bash
curl http://localhost:11434/api/chat -d '{"model": "llama70b",
  "messages": [{"role": "user", "content": "Hello!"}]}'
```

**Anthropic** — `/v1/messages`, `/v1/messages/count_tokens`

```bash
curl http://localhost:11434/v1/messages -H "Content-Type: application/json" \
  -d '{"model": "llama70b", "max_tokens": 256,
       "messages": [{"role": "user", "content": "Hello!"}]}'
```

A reasoning model's thinking never comes back as its answer. Each API gets it
where it puts reasoning: `message.thinking` on `/api/chat` (with `think`),
`reasoning_content` on `/v1/chat/completions`, a `reasoning` item on
`/v1/responses`, a `thinking` block on `/v1/messages` (with `thinking` enabled).
`think: false`, `reasoning_effort: "none"` and `thinking: {"type": "disabled"}`
tell the model not to think at all.

Models can be named by their local name, an alias, or the Hub reference they
were pulled from (`hf.co/org/repo:QUANT`). The server never downloads on its
own: `hfl pull` or `hfl serve --model <reference>` does.

## Reference

<details>
<summary><b>Configuration</b></summary>

| Variable | Default | What it does |
|---|---|---|
| `HFL_HOME` | `~/.hfl` | Where models, the registry and logs live |
| `HF_TOKEN` | — | Hugging Face token for gated models (or `hfl login`) |
| `HFL_MEMORY_BUDGET` | `85` | % of total RAM that may be in use after a load |
| `HFL_KEEP_ALIVE` | `5m` | How long an idle model stays loaded (`-1` = forever) |
| `HFL_MAX_LOADED_MODELS` | `0` | Optional ceiling on the number of loaded models |
| `HFL_LANG` | `en` | CLI language: `en` or `es` |

Settings Ollama also has (`OLLAMA_HOST`, `OLLAMA_KEEP_ALIVE`,
`OLLAMA_NUM_PARALLEL`, `OLLAMA_MAX_LOADED_MODELS`, …) are read under either
name. The full list is in [docs/env-vars.md](https://github.com/ggalancs/hfl/blob/main/docs/env-vars.md).

`hfl config` prints the directories, server address, rate limits, inference
defaults and timeouts in effect, environment variables applied.

Protect the API with a key: `hfl serve --api-key <secret>`, then send
`Authorization: Bearer <secret>` or `X-API-Key: <secret>`.

</details>

<details>
<summary><b>Concurrency and backpressure</b></summary>

llama.cpp and Transformers drive a single model instance that cannot take two
requests at once, so HFL runs inference one request at a time behind a bounded
queue, shared by the three APIs:

| Setting | Env var | Default |
|---|---|---|
| Requests running at once | `HFL_QUEUE_MAX_INFLIGHT` | `1` |
| Requests allowed to wait | `HFL_QUEUE_MAX_SIZE` | `16` |
| Seconds a request may wait | `HFL_QUEUE_ACQUIRE_TIMEOUT` | `60` |

A full queue answers **429** with `Retry-After`; a request that waited too long
answers **503**. Every response carries `X-Queue-Depth` and related headers, and
`GET /healthz` reports the live state.

</details>

<details>
<summary><b>Quantization levels</b></summary>

`Q4_K_M` is the default and the usual balance between size and quality. `Q5_K_M`,
`Q6_K` and `Q8_0` stay closer to the original model and take more memory; `Q3_K_M`
and `Q2_K` take less, at a visible cost in quality; `F16` is not quantized.
HFL tells you before loading whether a model fits — and `hfl recommend` suggests
the ones that do.

</details>

<details>
<summary><b>How it works</b></summary>

```text
hfl pull / run ──▶ Hugging Face Hub ──▶ ~/.hfl/models ──▶ GGUF? ── yes ──▶ llama.cpp
                   (search, download,                  MLX build (Apple Silicon) ──▶ MLX
                    license check)                     safetensors ── convert + quantize ──▶ GGUF

hfl serve ──▶ OpenAI · Ollama · Anthropic APIs ──▶ memory-budgeted model set ──▶ one inference at a time
```

The [architecture guide](https://htmlpreview.github.io/?https://github.com/ggalancs/hfl/blob/main/docs/hfl-architecture-complete.html)
covers the modules, engine selection, the conversion pipeline and every endpoint
([en español](https://htmlpreview.github.io/?https://github.com/ggalancs/hfl/blob/main/docs/hfl-arquitectura-completa.html)).

</details>

## Documentation

- [Hub-native features](https://github.com/ggalancs/hfl/blob/main/docs/hub-native-features.md) — discover, recommend, pull-smart, verify, bench and more
- [Environment variables](https://github.com/ggalancs/hfl/blob/main/docs/env-vars.md) — every setting and its default
- [Apple Silicon and Docker clients](https://github.com/ggalancs/hfl/blob/main/docs/apple-silicon-and-docker-clients.md)
- [Using HFL with other tools](https://github.com/ggalancs/hfl/blob/main/docs/integrations.md) — Open WebUI, AnythingLLM, Continue, aider, LangChain, LiteLLM: recipes run for real
- [Engine plugins](https://github.com/ggalancs/hfl/blob/main/docs/plugins.md) — add an engine from another package
- [Metrics](https://github.com/ggalancs/hfl/blob/main/docs/metrics.md) — what `/metrics` exports for Prometheus: speed, time to first token, memory, queue
- [Benchmarks](https://github.com/ggalancs/hfl/blob/main/docs/benchmarks.md) — HFL, Ollama and llama-server on the same GGUF, with the script to rerun it
- [Model compatibility](https://github.com/ggalancs/hfl/blob/main/docs/compatibility.md) — chat, tool calling, reasoning and vision checked for real on 13 families and both GGUF backends
- [Architecture guide](https://htmlpreview.github.io/?https://github.com/ggalancs/hfl/blob/main/docs/hfl-architecture-complete.html)
- [Running in production](https://github.com/ggalancs/hfl/blob/main/docs/production.md) — systemd, launchd, Docker Compose, Kubernetes, a TLS proxy, memory, monitoring: recipes run for real
- [Stability contract](https://github.com/ggalancs/hfl/blob/main/docs/stability.md) — what is public, and how it changes (deprecation first; SemVer from 1.0)
- [Changelog](https://github.com/ggalancs/hfl/blob/main/CHANGELOG.md)

**Status:** beta — 4,000+ tests at ~90% coverage. The full local audit
(`audit/`) runs on macOS, Linux and Windows 10; the Windows installers
(MSI, winget) are less tested than the rest.

## Contributing

Issues and pull requests are welcome — see [CONTRIBUTING.md](https://github.com/ggalancs/hfl/blob/main/CONTRIBUTING.md).

```bash
git clone https://github.com/ggalancs/hfl && cd hfl
pip install -e ".[dev]"
bash scripts/ci-local.sh        # lint, types and the full test suite, as CI runs them
```

If HFL saves you a download–convert–quantize afternoon, a ⭐ helps other people find it.

## Legal notices

**Model licenses.** Models keep their own licenses (Llama, Gemma, OpenRAIL,
CC-BY-NC, …) and you are responsible for complying with them. HFL shows a
model's license before downloading it, stores it with the model and records the
pull's provenance — see `hfl inspect <model>`. Common restrictions include
non-commercial use only (CC-BY-NC, MRL), attribution (Llama, Gemma) and usage
restrictions (OpenRAIL).

**Export compliance.** HFL only downloads publicly available open-weight models
from the Hugging Face Hub and does not facilitate access to closed-weight or
export-controlled weights. Users are responsible for complying with the export
regulations of their jurisdiction.

**Disclaimer.** AI models may generate inaccurate, biased or inappropriate
content. Users are solely responsible for evaluating and using model outputs.
See [DISCLAIMER.md](https://github.com/ggalancs/hfl/blob/main/DISCLAIMER.md).

**Trademarks.** "OpenAI" is a trademark of OpenAI, Inc. "Ollama" is a trademark
of Ollama, Inc. "Anthropic" is a trademark of Anthropic, PBC. "Hugging Face" and
the Hugging Face logo are trademarks of Hugging Face, Inc. These marks are used
for identification only. **HFL is an independent project, not affiliated with,
endorsed by or officially connected to any of these companies.** References to
their services describe technical interoperability only.

## License

HFL is licensed under the **Apache License 2.0** — you may use, modify,
distribute and sell it, including commercially, as long as you keep the
copyright and license notices. See [LICENSE](https://github.com/ggalancs/hfl/blob/main/LICENSE) and [NOTICE](https://github.com/ggalancs/hfl/blob/main/NOTICE).

HFL ships responsible-use safeguards: license checking, AI disclaimers,
provenance tracking, privacy protections and respect for gated models.
Apache-2.0 does not require you to keep them; as a project norm we ask that
redistributions leave them active. See [DISCLAIMER.md](https://github.com/ggalancs/hfl/blob/main/DISCLAIMER.md),
[PRIVACY.md](https://github.com/ggalancs/hfl/blob/main/PRIVACY.md) and [NOTICE-EU-AI-ACT.md](https://github.com/ggalancs/hfl/blob/main/NOTICE-EU-AI-ACT.md).

HFL's license covers HFL itself, not the models you download.
