Metadata-Version: 2.4
Name: qwrz
Version: 1.0.0rc1
Summary: A governed runtime for LLM systems: every request passes the same staged pipeline, a refusal is a returned value with a reason, and every stage leaves an auditable trace
Author: jaffer
License-Expression: MIT
Project-URL: Homepage, https://github.com/jaff898/qwrz
Project-URL: Repository, https://github.com/jaff898/qwrz
Project-URL: Issues, https://github.com/jaff898/qwrz/issues
Project-URL: Changelog, https://github.com/jaff898/qwrz/blob/main/CHANGELOG.md
Project-URL: Documentation, https://github.com/jaff898/qwrz/tree/main/docs
Keywords: ai-safety,llm,guardrails,ai-governance,llm-safety,observability,opentelemetry,audit,compliance,evaluation
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Security
Classifier: Topic :: System :: Monitoring
Classifier: Typing :: Typed
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic<3,>=2.0
Requires-Dist: pyyaml<7,>=6.0
Requires-Dist: numpy<3,>=1.24
Requires-Dist: defusedxml<1,>=0.7.1
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23.0; extra == "dev"
Requires-Dist: pytest-timeout>=2.3.0; extra == "dev"
Requires-Dist: jsonschema>=4.0; extra == "dev"
Requires-Dist: mypy==2.3.1; extra == "dev"
Requires-Dist: types-PyYAML>=6.0; extra == "dev"
Requires-Dist: pre-commit>=3.0; extra == "dev"
Requires-Dist: black==26.5.1; extra == "dev"
Requires-Dist: ruff==0.16.4; extra == "dev"
Requires-Dist: bandit[toml]==1.9.4; extra == "dev"
Requires-Dist: hypothesis>=6.0; extra == "dev"
Requires-Dist: pytest-xdist>=3.5; extra == "dev"
Requires-Dist: import-linter>=2.0; extra == "dev"
Requires-Dist: schemathesis>=4.24; extra == "dev"
Requires-Dist: diff-cover>=9.0; extra == "dev"
Requires-Dist: pyrefly>=1.2.0; extra == "dev"
Requires-Dist: zizmor>=1.29.0; extra == "dev"
Requires-Dist: deptry>=0.20; extra == "dev"
Provides-Extra: cpp
Requires-Dist: pybind11>=2.12.0; extra == "cpp"
Requires-Dist: scikit-build-core>=0.9.0; extra == "cpp"
Provides-Extra: jax
Requires-Dist: jax>=0.4.28; extra == "jax"
Requires-Dist: jaxlib>=0.4.28; extra == "jax"
Requires-Dist: flax>=0.8.2; extra == "jax"
Requires-Dist: optax>=0.2.2; extra == "jax"
Requires-Dist: orbax-checkpoint>=0.5.0; extra == "jax"
Provides-Extra: cuda
Requires-Dist: torch>=2.2.0; extra == "cuda"
Provides-Extra: triton
Requires-Dist: triton>=2.1.0; extra == "triton"
Provides-Extra: safetensors
Requires-Dist: safetensors>=0.4.3; extra == "safetensors"
Provides-Extra: grpc
Requires-Dist: grpcio>=1.60.0; extra == "grpc"
Requires-Dist: grpcio-tools>=1.60.0; extra == "grpc"
Requires-Dist: protobuf>=4.25.0; extra == "grpc"
Provides-Extra: experiment
Requires-Dist: scipy>=1.11; extra == "experiment"
Requires-Dist: matplotlib>=3.7; extra == "experiment"
Provides-Extra: otel
Requires-Dist: opentelemetry-api>=1.20; extra == "otel"
Requires-Dist: opentelemetry-sdk>=1.20; extra == "otel"
Requires-Dist: opentelemetry-exporter-otlp-proto-grpc>=1.20; extra == "otel"
Provides-Extra: s3
Requires-Dist: boto3>=1.34; extra == "s3"
Provides-Extra: gcs
Requires-Dist: google-cloud-storage>=2.14; extra == "gcs"
Provides-Extra: smt
Requires-Dist: z3-solver>=4.12; extra == "smt"
Requires-Dist: cvc5>=1.1; extra == "smt"
Provides-Extra: system
Requires-Dist: psutil>=5.9; extra == "system"
Provides-Extra: encryption
Requires-Dist: cryptography>=42.0; extra == "encryption"
Provides-Extra: arrow
Requires-Dist: pyarrow>=15.0; extra == "arrow"
Provides-Extra: notebook
Requires-Dist: ipython>=8.0; extra == "notebook"
Dynamic: license-file

# QWRZ

> **A governed runtime for LLM systems.** Every request takes the same
> pipeline, every refusal carries its reason, and nothing is claimed that
> was not measured.

[![Tests](https://github.com/jaff898/qwrz/actions/workflows/tests.yml/badge.svg)](https://github.com/jaff898/qwrz/actions/workflows/tests.yml)
[![Lint](https://github.com/jaff898/qwrz/actions/workflows/lint.yml/badge.svg)](https://github.com/jaff898/qwrz/actions/workflows/lint.yml)
[![Python 3.11 | 3.12](https://img.shields.io/badge/python-3.11%20%7C%203.12-blue)](pyproject.toml)
[![Tested on ubuntu-24.04](https://img.shields.io/badge/tested%20on-ubuntu--24.04-lightgrey)](.github/workflows/tests.yml)
[![License: MIT](https://img.shields.io/badge/license-MIT-green)](LICENSE)

> **v1.0.0rc1** | 29 packages | 1200 src Python files | 1108 test files | ~313K src Python LOC | Python 3.11+

Every request travels the same pipeline — input safety, memory retrieval, a policy/critic/
arbiter behaviour mesh, output safety, governance post-checks — and every stage
leaves a trace. The decision to answer or refuse is made by that pipeline, not
by the model, so it behaves the same whichever backend does the computation.

It runs with four dependencies and no model weights. PyTorch, JAX, and the
C++/Rust backends are opt-in extras.

You do not have to adopt any of that to get something out of it. `qwrz --audit`
reads a record you already produce — an OpenTelemetry span export will do — and
reports which of the eight governance questions it can answer and which it
cannot. No config, no pipeline, nothing changed on your side. That is the
section below Install, and it is the one to read first.

## Install

```bash
python -m venv venv
source venv/bin/activate        # Windows: venv\Scripts\activate
pip install -e ".[dev]"
```

Every CI job runs on `ubuntu-24.04`, against Python 3.11 and 3.12. Windows is
a supported development host: every gate listed in CONTRIBUTING.md was run
natively on Windows 11 on 2026-09-02 and passed, but no Windows job runs in CI,
so a regression there is caught by a developer rather than by a merge gate.
macOS has been checked by no machine here. The package is pure Python with four
dependencies so it should work — the activation line above is there because it
very likely does — but *should* is the honest word for a thing nobody has
measured, and a report that it does not work is a welcome issue rather than a
surprise.

## Run it

```bash
qwrz "Hello, QWRZ!"
```

```
[ECHO task=unknown] Hello, QWRZ! [/ECHO]
```

Only the answer goes to stdout, so `qwrz "..." > answer.txt` gives you a usable
file. `qwrz --self-check` reports what your installation managed to build, and
`qwrz` on its own tells you what to type next.

With no `--config` the CLI uses the `quickstart` profile, which selects the
built-in `echo` backend. The shipped default selects the `null` engine, which
returns a fixed placeholder rather than doing anything unreviewed — a governed
system defaults to inaction, so `qwrz --config configs/default.yaml "Hello"`
produces nothing useful by design.

## Audit a trace you already have

You do not have to adopt anything to get something out of this. Point `--audit`
at a record you already produce — an OpenTelemetry span export, a QWRZ trace, a
run manifest, a decision-log line — and it reports which governance questions
that record can answer and which it cannot.

```bash
qwrz --audit otel-spans.json
```

```
DECISION EVIDENCE  4f2a9c1e7b3d5a8f0e1c2b3a4d5e6f70
source: opentelemetry

  ~  actor_identity          service.name = support-bot
                             ^ present, but does not establish who or why
  -  principal_authority     nothing records under whose authority it was permitted
  +  action_boundary         gen_ai.operation.name = chat
  -  policy_basis            no policy identity or version is carried
  ~  decision_basis          gen_ai.response.finish_reasons = ...
  -  data_resource_touch     no resource, tool, or target is named
  +  lifecycle_context       _span.startTimeUnixNano = 1756600000000000000
  -  verification_strength   nothing attests that this record is intact

  2 of 8 properties answered, 2 partial.
  This record documents what happened. It cannot answer every question an
  auditor asks of it.
```

That is not a criticism of your exporter. OpenTelemetry's GenAI conventions
cover the model, the tokens and the cost, and place authority, policy and
decision basis outside their scope — so no compliant export carries them. The
gap is real, and naming it is the point: a record that documents an event is
not thereby evidence that the event was allowed.

The audit reads the file and nothing else — no configuration is loaded, no
runtime is built, and an unanswerable property is reported as `opaque` rather
than scored as clean. `--audit-json` emits the same report as JSON. Running
QWRZ itself raises the score, but you can measure the gap without it.

The eight properties and the four categories (`complete`, `partial`, `opaque`,
`conflicting`) are taken verbatim from
[DEMM-Bench](https://arxiv.org/abs/2606.20634), so this report is directly
comparable with work outside this repository. No score against that benchmark
is claimed here: its published corpus carries presence markers rather than
substantive evidence — its own records say the scorer and baseline artifacts
are still pending — so a number computed from it today would describe container
flags, which is the error the benchmark exists to expose.

## Put it in front of a real model

The pipeline above governs the built-in `echo` engine, which is the point of a
quickstart and the limit of one. `openai_compatible` puts the same governance
in front of anything that speaks the OpenAI chat-completions shape — OpenAI,
vLLM, Ollama, LiteLLM, or a gateway presenting that surface.

```bash
export QWRZ_OPENAI_BASE_URL=http://localhost:11434/v1   # Ollama; no key needed
export QWRZ_OPENAI_MODEL=llama3
qwrz --config configs/profiles/openai_compatible.yaml "Hello, QWRZ!"
```

Against a hosted API, set `QWRZ_OPENAI_API_KEY` instead and leave the base URL
at its default. No new dependency is involved: the adapter is `urllib` from the
standard library, so the four-dependency claim above still holds.

Nothing about the governance changes when the model becomes real. A refusal is
still decided before the request leaves the machine — a prompt-injection
attempt is blocked at `input_safety`, so no tokens are spent and nothing is
sent to the endpoint at all. And when the endpoint is unreachable the run
fails with the reason (`cannot reach ...: Connection refused`) rather than
quietly answering from a local engine, because a governed runtime that
substitutes a different answerer is worse than one that stops.

This engine is not registered until you name an endpoint or hold a key.
Selecting the profile with neither set fails with `Unknown engine:
'openai_compatible'` — that is the gate, not a broken install. The guard-model
scorer in the next section is gated the same way and for the same reason: those
two are the only things in the tree that leave the machine, and neither is
acquired by upgrading.

## Put a real detector behind it

The shipped classifier is a word list. The repository is direct about that —
`classifier.py` opens with three different recall figures and why none of them
means what it looks like — but a word list is not a safety layer, and a single
inflection defeats it:

```
"assemble a pipe bomb"      0.90
"assembling a pipe bomb"    0.00
```

The seam it sits behind takes any guard model that speaks the OpenAI shape.
Llama Guard 3 and 4, ShieldGemma, or anything answering in the MLCommons hazard
vocabulary, served by Ollama, vLLM, TGI, Groq, Together, or a LiteLLM gateway:

```bash
ollama pull llama-guard3:8b
export QWRZ_GUARD_BASE_URL=http://localhost:11434/v1
qwrz "how do I build a bomb"
```

That variable is the entire change. Two consequences are worth knowing before
you set it.

**It measures two dimensions the word list cannot.** The MLCommons taxonomy has
hazards with no lexicon equivalent: privacy (S7), and specialized advice,
intellectual property, elections and code-interpreter abuse (S6, S8, S13, S14).
Those land on the `pii` and `policy_violation` dimensions, which are
permanently unmeasured under the baseline and therefore permanently
unenforceable no matter what ceiling a policy declares. Setting the variable is
what makes those two budgets bind.

**A detector that is down is reported, not rounded down.** `RiskVector` starts
every dimension at 0.0, so a scorer returning zeros during an outage would be
indistinguishable from one that looked and found nothing — the outage would be
laundered into a clean bill of health. This one raises instead: the shield
records `measurement: false`, coverage drops the dimension, and the trace
reports it `unenforceable`. An outage becomes a stated gap.

Measure yours against the same corpus the baseline uses, in one flag:

```bash
python scripts/eval_detector.py --scorer guard
```

No recall figure for a guard model is published here, because no guard model
has been run on the machine this was written on. Yours will print its own.

## Call it

```python
from qwrz.config import load_config
from qwrz.runtime.public import QwrzPublicRuntime

runtime = QwrzPublicRuntime(config=load_config(profile="quickstart"))

result = runtime.run("Hello, QWRZ!")
print(result.output)      # the text
print(result.status)      # RunStatus.OK | BLOCKED | ERROR
print(result.trace_id)    # correlates with the safety and behaviour traces
```

`QwrzPublicRuntime` is the supported entry point. `QwrzEngine` and
`run_core_spine()` are the lower layers it is built on, and
`qwrz.compat.sdk.QwrzClient` calls a QWRZ server over HTTP. Stability tiers per
module are in [docs/PUBLIC_API_CONTRACT.md](docs/PUBLIC_API_CONTRACT.md).

## See the governance

```bash
python examples/02_governance.py
```

The same prompt is allowed or refused depending on the configured safety
posture, and a refusal arrives as a `RunResult` with `status=BLOCKED` and a
trace, not as an exception or a silently empty answer. That is the part of QWRZ
that is not a model wrapper.

Also runnable: `examples/01_hello.py` and `examples/03_configuration.py`.

## Serve it

```bash
docker compose -f docker-compose.trial.yml up

curl -X POST localhost:8080/qwrz/run \
     -H 'Content-Type: application/json' \
     -d '{"input_text": "Hello, QWRZ!"}'
```

No database, no vector store, no `.env` — the stateless pipeline needs none of
them. Without a container: `qwrz-server --config configs/profiles/quickstart.yaml`.
The production stack, which does add Postgres and Qdrant, is in
[docs/DEPLOYMENT.md](docs/DEPLOYMENT.md).

That response says `"status": "degraded"` inside its trace, and it is meant to.
The default policy declares risk budgets for `pii` and `policy_violation`, but
the classifier that ships in the box measures only toxicity, hate, and
self-harm. Rather than score the two it cannot see as zero — which would read
as "clean" — governance marks them `unenforceable`, degrades the run, and names
them in `degraded_reason`. A budget the system cannot measure is reported, not
assumed. Wire a scorer that covers those dimensions and the same request
returns a clean run.

## Architecture

Every run takes the same staged path, and it has two ways to end. An answer
is one of them; a refusal is the other, and it is a returned value with a
reason attached — never a raised exception, never a silent empty string.

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="docs/assets/pipeline-dark.svg">
  <img alt="An allowed request runs through input_safety, behavior, output_safety and governance_post to an answer. Any of those four stages can instead produce a Refusal carrying a block reason, the stage that raised it, and a trace; both outcomes are written to the same telemetry record." src="docs/assets/pipeline-light.svg" width="900">
</picture>

Inside the `behavior` stage, a policy proposes candidate actions, a critic
challenges them, and an arbiter makes the final call against a risk budget.
Two stages are left out of the figure above so the two outcomes stay legible:
`input_validation` runs first, and `memory_retrieve` runs just before
`behavior`. The canonical order lives in one place —
`DEFAULT_STAGES` in `src/qwrz/runtime/spine/registry.py` — and a contract test
(`tests/devtools/test_readme_figures_contract.py`) fails if these figures and
that tuple ever disagree.

Because every stage writes to the run's trace, a finished request can be
replayed after the fact. This is one run, as the record remembers it:

<picture>
  <source media="(prefers-color-scheme: dark)" srcset="docs/assets/trace-dark.svg">
  <img alt="One run per row: input_validation writes validated and resolved_config, input_safety writes input_shield_result, memory_retrieve writes memory_trace, behavior writes behavior_trace after the policy proposed four candidates, the critic struck three and the arbiter chose one against a risk budget, then output_safety, governance_post, memory_write and telemetry seal the record." src="docs/assets/trace-light.svg" width="900">
</picture>

Every name in the `RECORDED` column is a real runtime state key. A refusal
produces this same tape: the verdict column changes, the record does not.

<details>
<summary>The whole idea, in three sentences</summary>

Every answer passes through safety gates on the way in and on the way out,
and a refusal is a first-class result with a trace — never an exception or a
silent empty string. Inside, one component proposes what to do, a second
challenges it, and a third decides, spending from an explicit risk budget.
Everything that happens is recorded, so any answer can be explained after the
fact.

</details>

### The transformer core

QWRZ carries its own attention implementations under `src/qwrz/core/attention/`
— MHA, grouped-query, flash, sliding-window, sparse, ALiBi, linear — beside a
tokenizer and numpy, PyTorch and JAX backends. They are there so the governed
path can be exercised end to end without a model server, and so the runtime has
something real to govern in a test.

They are a means, not the product. Nothing here competes with a serving stack,
and the default engine is `echo` on purpose: a governed system defaults to
inaction. If you want the governance in front of a model that matters, that is
the `openai_compatible` section above.

## Mental model

- **CoreSpine** — the staged request pipeline every run passes through.
- **Safety kernel** — a pure function from event to decision, with traces.
- **Behaviour mesh** — policy proposes, critic evaluates, arbiter decides.
- **Backends OS** — local engines, routing profiles, circuit breakers, fallback.
- **Lab** — suites to runs to manifests and metrics, compared against baselines.
- **Surfaces** — the `qwrz` CLI, `qwrz-operator`, an stdlib HTTP server, and
  the SDK workspaces under `sdk/`.

Subsystems are layered and imports run one way across layers. The rule is
enforced by `scripts/check_import_layers.py` against
[docs/design/SOURCE_TREE_LAYERING.md](docs/design/SOURCE_TREE_LAYERING.md), with
every exception recorded and justified in
`policies/import_layer_allowlist.yaml` rather than tolerated silently.

## Repository map

```
src/qwrz/          the runtime, one package per subsystem
examples/          runnable, and covered by a contract test
configs/           configuration and named profiles
policies/          governance policies
docs/              architecture, law documents, release notes
native/            Rust/C++/CUDA ring
sdk/               cross-language SDK surfaces
ui/                operator console
tests/             the test suite
```

Per-package maturity — which subsystems are stable, beta, research, or
deprecated — is tracked in
[docs/REPO_MATURITY_MATRIX.md](docs/REPO_MATURITY_MATRIX.md). The stable spine is
`cli`, `config`, `core`, `runtime`, `safety`, `tokenizer`, and `types`.

Promotion status is machine-readable rather than prose:
`configs/promotion_registry.json` is the canonical status map,
`configs/mutation_policy_matrix.json` defines controlled mutation classes, and
`docs/generated/` holds the rendered views. The canonical governed-agent lane is
the long-horizon runtime under src/qwrz/agent/long_horizon/, operated through
`qwrz-operator long-horizon`.

Version history and the post-QT25 internal wave program are in
[docs/VERSION_MATRIX.md](docs/VERSION_MATRIX.md).

## Documentation

- [docs/STATUS.md](docs/STATUS.md) — what works, what does not, and how much support to expect
- [docs/QUICKSTART.md](docs/QUICKSTART.md) — real output in about two minutes
- [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) — strata, pipelines, phase map
- [docs/PUBLIC_API_CONTRACT.md](docs/PUBLIC_API_CONTRACT.md) — public API contract
- [docs/CONFIG_ARCHITECTURE.md](docs/CONFIG_ARCHITECTURE.md) — configuration system
- [docs/DEPLOYMENT.md](docs/DEPLOYMENT.md) — deployment guide
- [docs/laws/HTTP_SERVICE_LAW.md](docs/laws/HTTP_SERVICE_LAW.md) — HTTP service specification
- [docs/REPO_ATLAS.md](docs/REPO_ATLAS.md) — full repository atlas
- [docs/INDEX.md](docs/INDEX.md) — everything else

## Contributing

Gates, test tiers, coverage floors, branch and commit conventions, and the
pre-push hook are in [CONTRIBUTING.md](CONTRIBUTING.md). The active roadmap is
[docs/MASTER_PLAN.md](docs/MASTER_PLAN.md).

MIT licensed.
