Metadata-Version: 2.4
Name: agentic-security-harness
Version: 1.9.0
Summary: Open-source defensive harness and learning lab for reproducing and measuring agentic AI failure modes with portable traces, attack graphs, scorecards, and data-boundary tests.
Author: Dmitry Krivonosov
License-Expression: Apache-2.0
Project-URL: Homepage, https://github.com/krivonosoff161/agentic-security-harness
Project-URL: Repository, https://github.com/krivonosoff161/agentic-security-harness
Project-URL: Issues, https://github.com/krivonosoff161/agentic-security-harness/issues
Project-URL: Changelog, https://github.com/krivonosoff161/agentic-security-harness/blob/main/CHANGELOG.md
Keywords: ai-security,agentic-ai,benchmark,llm-security,prompt-injection,red-team
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Information Technology
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Security
Classifier: Typing :: Typed
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: pydantic<3,>=2
Provides-Extra: transfer
Requires-Dist: agentic-transfer-verifier==0.2.1; extra == "transfer"
Requires-Dist: agentic-transfer-verifier-harness-extension==1.0.1; extra == "transfer"
Provides-Extra: handoff
Requires-Dist: ai-agent-handoff==0.3.0; extra == "handoff"
Requires-Dist: ai-agent-handoff-harness-extension==1.0.0; extra == "handoff"
Provides-Extra: playbooks
Requires-Dist: llm-safety-playbooks==0.1.0; extra == "playbooks"
Provides-Extra: router
Requires-Dist: agentic-llm-router==0.2.1; extra == "router"
Provides-Extra: filter
Requires-Dist: llm-cheap-filter==0.2.0; extra == "filter"
Provides-Extra: all
Requires-Dist: agentic-transfer-verifier==0.2.1; extra == "all"
Requires-Dist: agentic-transfer-verifier-harness-extension==1.0.1; extra == "all"
Requires-Dist: ai-agent-handoff==0.3.0; extra == "all"
Requires-Dist: ai-agent-handoff-harness-extension==1.0.0; extra == "all"
Requires-Dist: llm-safety-playbooks==0.1.0; extra == "all"
Requires-Dist: agentic-llm-router==0.2.1; extra == "all"
Requires-Dist: llm-cheap-filter==0.2.0; extra == "all"
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: ruff>=0.5; extra == "dev"
Requires-Dist: mypy>=1.8; extra == "dev"
Dynamic: license-file

# Agentic Security Harness

[![OpenSSF Best Practices](https://www.bestpractices.dev/projects/13320/badge)](https://www.bestpractices.dev/projects/13320)
[![CI](https://github.com/krivonosoff161/agentic-security-harness/actions/workflows/ci.yml/badge.svg)](https://github.com/krivonosoff161/agentic-security-harness/actions/workflows/ci.yml)
[![CodeQL](https://github.com/krivonosoff161/agentic-security-harness/actions/workflows/codeql.yml/badge.svg)](https://github.com/krivonosoff161/agentic-security-harness/actions/workflows/codeql.yml)
![Python](https://img.shields.io/badge/python-3.11%2B-blue)
![License](https://img.shields.io/badge/license-Apache--2.0-green)
![Status](https://img.shields.io/badge/public_research_release-v1.8.0-blue)

**Your AI coding agent reads untrusted repository text. Can it keep data separate from
instructions and authority?**

Agentic Security Harness is a local, trace-first benchmark for defensive testing of
agentic AI boundary failures. It runs reproducible synthetic scenarios, records portable
traces and scorecards, and compares a deliberately vulnerable local agent with a protected
one.

In plain English: it turns “the agent behaved unsafely” into evidence you can replay,
validate, compare, and review.

## In development — controlled file effects

[Issue #326](https://github.com/krivonosoff161/agentic-security-harness/issues/326)
tracks a [1.9.0 candidate](docs/controlled-file-workflow.md) with actual writes to
fresh synthetic files, report-only Guard enforcement and independent disk checks.
It is not in the published 1.8.0 package. Its [fresh eight-call model observation](docs/controlled-file-observation-20261001.md)
recorded four forbidden writes denied, four allowed reports written and two correct
reports. Release checks are separate; the historical observations below are unchanged.

## Latest observation — 2026-09-27

The [sixteen-call generalization study](docs/proposal-generalization-20260927.md)
records **four exact useful tasks**, one permitted but semantically wrong hash
operation, three rejected useful-task proposals and eight negative controls stopped
at their declared boundaries. Five pure operations are not five correct tasks.
The local Prometheus model was real; downstream Router transport and adversarial
receipt mutations were declared fixtures. Real effects were zero; there was no
response repair or retry. This is separate from the garak package-release gates.

## Previous verified result — 2026-09-26

The [twelve-call local-model follow-up](docs/ollama-quarantine-adapter.md#proposal-contract-follow-up-2026-09-26)
records **three real-model proposals completing all seven boundaries** of the
installed ecosystem, each ending in a built-in constant lookup. All six negative
controls stopped at their declared boundary; real external effects were zero.
Explicit protocol wording matched both fixed tasks; free wording matched neither
capability identifier. This is bounded integration evidence, not a model reliability
estimate, an autonomous-agent demonstration or an independent human audit.

- **Inspect:** [fixed public corpus](examples/installed-ecosystem/proposal-contract-cases.json),
  [content-free observations](examples/installed-ecosystem/proposal-contract-observation.json)
  and [artifact checks](tests/test_proposal_contract_evidence.py).
- **Use the example:** [explicit proposal protocol](examples/installed-ecosystem/README.md#explicit-proposal-protocol-and-real-model-follow-up).
  Printing the prompt or checking artifacts does not rerun the model.
- **Delivery:** [merged PR #303](https://github.com/krivonosoff161/agentic-security-harness/pull/303).
  The earlier eight-call result is preserved separately in the report.

**Package versus research evidence:** the study used exact installed **1.6.0**
subjects, while the current published package is **1.8.0**. The prompt helper,
corpus and observations are repository-owned evidence, not evidence that the
1.7.0, 1.7.1 or 1.8.0 wheel was used in those model calls. See
[release and package status](#release-and-package-status) and
[current state](docs/current-state.md) for the exact boundary.

## Experimental garak connection

The published **1.7.0** wheel packages this opt-in adapter. Its attested tag
build and exact-wheel TestPyPI staging passed on Linux and Windows; the PyPI
upload matches those subjects. Initial Python 3.12/3.13 smokes stopped at a
missing same-version PyYAML wheel hash before garak. The
[corrected read-only verification](https://github.com/krivonosoff161/agentic-security-harness/actions/runs/36306894913)
passed all seven jobs without a new upload. [Release evidence and gates](docs/releases/v1.7.0.md) keep
the prior 1.6.0 evidence separate.

The [garak plan connector](docs/garak-plan-connector.md) translates one strictly
validated JSON plan into the existing Quarantine/Gateway path, without granting
authority from a detector score. A [four-case example](examples/garak-gateway/README.md)
keeps detection, admission, policy denial and synthetic execution separate. This is
an independent Harness integration, not an official NVIDIA/garak component. It is
new in the 1.7.0 package, **not included in published 1.6.0**; its optional detector
compatibility lane pins an unmerged upstream PR rather than claiming release support.

## Retained ancestry — published in 1.8.0

[PR #324](https://github.com/krivonosoff161/agentic-security-harness/pull/324)
adds an opt-in [retained ancestry store](docs/ancestry-store.md): trusted root
admission, exact retained parent history, scope narrowing and local crash recovery.
Ancestry acceptance cannot override Gateway policy. The exact release subjects are
on GitHub, TestPyPI and PyPI; a separate read-only run passed all seven cross-index
jobs after the initial production Linux 3.11 index lookup failed before application
checks. [Release status](docs/releases/v1.8.0.md) binds the source, hashes and limits.

## Quickstart

Published [v1.8.0](docs/releases/v1.8.0.md) adds opt-in retained ancestry admission.
The prior [v1.7.1 patch](docs/releases/v1.7.1.md#publication-evidence) repairs
known-empty parent authority scope and memory TTL narrowing. The opt-in garak
plan adapter, native Ollama proposal adapter, installed-ecosystem pilot,
Gateway lookup-key type guard and Router 0.2.1 remain. Release notes distinguish
package availability, retained initial smoke failures, and successful read-only
verification.

Install the exact package version from
[PyPI](https://pypi.org/project/agentic-security-harness/1.8.0/):

```bash
python -m pip install agentic-security-harness==1.8.0
ash quickstart --out reports/quickstart
ash agent-host-quickstart --out reports/agent-host-quickstart
```

For source development:

```bash
git clone https://github.com/krivonosoff161/agentic-security-harness.git
cd agentic-security-harness
python -m pip install .
ash quickstart --out reports/quickstart
ash agent-host-quickstart --out reports/agent-host-quickstart
```

`ash quickstart` is local and no-network. It runs the same stable 24-pattern corpus against
both demo targets, validates the generated artifacts, and renders a self-contained HTML
report.

`ash agent-host-quickstart` is the shipped provider-neutral integration path in v1.1.0.
It runs a built-in owned synthetic workflow through the
instrumented collector, canonical recordings, deterministic evaluator, atomic bundle
publication, and shared validator. It retains digest-only public evidence and makes no
provider call.

### Runtime Gateway synthetic contour

The source tree now also contains a runnable local Runtime Gateway increment. It applies a
closed policy before synthetic tool dispatch, exposes bounded OpenAI-compatible and MCP
2026-07-28 stateless development endpoints, and maintains a privacy-minimized append-only
audit chain:

```bash
ash gateway-init --out gateway.toml
ash gateway-check --config gateway.toml
ash gateway-serve --config gateway.toml
```

Open <http://127.0.0.1:8787/dashboard> after startup, or use the hardened Docker Compose
profile in [Runtime Gateway synthetic contour](docs/runtime-gateway.md). This is a
credential-free synthetic integration surface. Offline
[provider-neutral tool-call adapters](docs/provider-tool-adapters.md) normalize retained
OpenAI Responses, Anthropic Messages, Google Interactions, and MCP payloads through the
same policy without SDKs or credentials. Live provider transport and a production firewall
are still not shipped. The gateway exposes its exact closed policy and non-executable
approval-request digests; it intentionally has no approval-grant endpoint yet.

### Ecosystem map

Harness is the released core and the public contract owner for a modular security
ecosystem. Transfer verification, handoff safety, playbooks, routing, filtering, private
Runtime Guard research, and the public profile keep their own source-owned component
facts. The Harness generates only the cross-project roadmap and compatibility view:

- [Ecosystem roadmap](docs/ecosystem-roadmap.md)
- [Components and current integration status](docs/ecosystem-components.md)
- [Documentation crosswalk](docs/documentation-map.md)
- [`component.yaml`](component.yaml) and [`ecosystem/roadmap.yaml`](ecosystem/roadmap.yaml)

Runtime Guard remains private and `contract_only`. Harness version `v1.8.0`
contains the closed [Extension SDK V1](docs/extension-sdk.md) and public passive extras for
validated observation-to-finding dataflow. It does not auto-load installed packages;
companion repositories remain optional, separately versioned distributions.

Version `v1.8.0` retains the closed optional-dependency groups introduced in `v1.4.0`:

| Extra | Exact companion distributions | Automatic activation |
|---|---|---|
| `transfer` | `agentic-transfer-verifier==0.2.1`, extension `==1.0.1` | no |
| `handoff` | `ai-agent-handoff==0.3.0`, extension `==1.0.0` | no |
| `playbooks` | `llm-safety-playbooks==0.1.0` data-only wheel | no |
| `router` | `agentic-llm-router==0.2.1` | no |
| `filter` | `llm-cheap-filter==0.2.0` | no |
| `all` | the exact union of the five rows | no |

The generic PyPI coordinate `llm-router` is intentionally absent because it belongs to
another project. CI builds all eight exact wheels from pinned Git SHAs and installs the
closed local wheelhouse without loading either extension entry point. The public install
commands are
`pip install "agentic-security-harness[router]==1.8.0"` or
`pip install "agentic-security-harness[all]==1.8.0"`. Other companion pins are unchanged;
installation remains separate from module activation.

For explicit installed extension binding, follow the fresh `--no-compile` environment
and positive/negative controls in the
[installed-ecosystem example](examples/installed-ecosystem/README.md). The
[external pilot protocol](docs/external-pilot.md) explains the bounded feedback request.

For the actual six-component flow, run the separate
[16-case functional chain](examples/installed-ecosystem/README.md#functional-six-component-chain)
and its independent report verifier. Router and Filter APIs execute here; provider
transport is an explicit in-memory double. The [expert-readiness contract](docs/expert-readiness.md)
separates this reproducible technical baseline from outstanding human review.

Published Router 0.2.1 includes the lazy HTTP-client import repair; Harness binds
its reviewed receipt source to that exact release commit. Installing the package does
not call a provider or make the receipt an action grant.

The published [Corpus Pack SDK V1](docs/corpus-pack-sdk.md) adds a separate,
canonical registry for optional namespaced boundary-invariant metadata. It preserves the
frozen corpus 1.0.0, loads no package code, and treats complete evidence as readiness for
later rule evaluation rather than a security verdict.

The published v1.3.0 [companion adapter contracts](docs/companion-extensions.md) exercise
exact Transfer Verifier reports, Handoff metadata and Playbooks guidance through that
SDK on Linux and Windows. This closes a concrete producer-to-consumer dataflow gap; it
does not auto-load those installed distributions or make them production enforcement.

The published v1.3.0 [Security Intelligence contour](docs/security-intelligence-extension.md)
adds a provider-neutral offline weekly public-source review contract with digest-only
evidence, explicit coverage gaps, and no live fetching or model authority.

The published v1.3.0 [receipt auditors](docs/receipt-auditor-extensions.md)
independently checks exact-pinned Router invocation and Cheap Filter triage accounting
receipts from caller-supplied canonical bytes. Valid accounting remains `inconclusive`,
missing evidence remains `inconclusive`, and drift becomes a finding; the auditors never
emit `pass`, import or invoke the companion packages, or lower a security decision.

The published [Extension Distribution Discovery V1](docs/extension-distribution-discovery.md)
inspects one explicitly named local distribution without importing it. It verifies its
`RECORD`, closed entry point, canonical manifest, implementation bytes and caller-supplied
configuration digest, then requires exact reinspection before issuing an authority-free
approval receipt. Harness still does not load package code: the operator supplies an
already constructed object, and the binder checks it against the approved manifest pins.

The published [Extension Operator Lifecycle V1](docs/extension-operator-lifecycle.md)
exposes that metadata-only inspection and exact-reinspection approval through safe CLI
commands, adds canonical disable and non-executable rollback-plan receipts, and lists only
explicitly supplied receipt state. It never imports, downloads, starts, stops, or rolls
back extension code; the embedding application must construct and bind an object and
enforce any accepted disable artifact.

The published [controlled local adapter](docs/controlled-local-adapter.md)
connects only to an operator-started literal-loopback `/v1/responses` endpoint and passes
canonical tool calls through the existing closed Runtime Gateway policy. It supports local
model names as opaque identifiers—including Qwen and DeepSeek-style names—without vendor
claims. It has no DNS, proxy, redirect, credential, external-provider, arbitrary-tool, or
upstream-MCP path; receipts are digest-only and operational authority remains `none`.

The published [Policy Pack extension](docs/policy-pack-extension.md) independently
parses one exact-pinned data-only Playbooks pack and evaluates caller-supplied content-free
signals bound to canonical observations. A missing pack is inconclusive; production
Harness does not import or execute Playbooks code, discover packages, call a network, or
grant allow/enforcement authority.

| Target | Modeled findings | Patterns passed |
|---|---:|---:|
| `demo-agent` | 24 | 0 |
| `protected-demo-agent` | 0 | 24 |

The deterministic result is **24 modeled findings** for the vulnerable fixture and
**0 modeled findings** for the protected fixture. This is synthetic conformance evidence,
not a production safety claim.

![Terminal comparison showing 24 findings reduced to 0](docs/images/terminal-compare.png)

![Rendered comparison report table](docs/images/comparison-table.png)

Inspect the committed before/after example in
[`examples/comparison-report/`](examples/comparison-report/) or validate every public
example locally:

```bash
ash validate examples/
ash validate docs/evidence-status-registry.json
```

## What this is / is not

| This project is | This project is not |
|---|---|
| A reproducible benchmark for agent operating-environment boundaries. | A production safety certification. |
| A synthetic and authorized defensive testing lab. | A live exploitation or persistence toolkit. |
| A way to compare vulnerable and protected targets using portable artifacts. | Proof that a provider, model, or deployed agent is secure. |
| A stable trace/corpus contract with machine-readable validation. | A model leaderboard or CVE database. |

The benchmark focuses on agent operating-environment boundaries, not just standalone model answers.
Built-in targets are deterministic and offline. The experimental external adapter
is explicit opt-in, prompt-only, and does not execute tools. See
[`docs/benchmark-semantics.md`](docs/benchmark-semantics.md) and
[`docs/authorized-testing-paths.md`](docs/authorized-testing-paths.md).

## Visual evidence snapshot

![Evidence flow from scenario to validated report](docs/assets/evidence-flow.svg)

The public evidence map separates deterministic executable specifications, sanitized local
observations, historical material, and independently reviewed evidence:

- [Evidence map](docs/showcase/evidence-map.md)
- [Evidence classes](docs/evidence-classes.md)
- [Machine-readable evidence registry](docs/evidence-status-registry.json)
- [Evidence pack format](docs/evidence-pack-format.md)
- [Private/public evidence boundary](docs/private-public-evidence-boundary.md)

Public artifacts may include scenario identifiers, aggregate counts, response hashes, and
validator results. They do not include raw private prompts, raw responses, synthetic
canaries, or local paths. Keep external raw material under
`.internal/external-demo/latest`. **Do not** commit `raw_responses/` or other private
runtime evidence.

## Current stable surface

- Trace schema `1.0`, with a bounded legacy `0.1` read/migration window.
- Corpus `1.0.0`, freezing 24 ordered synthetic pattern identifiers.
- Local targets for vulnerable/protected agent, RAG, tool, function, and multi-agent
  handoff comparisons, including `toy-multi-agent` and `protected-toy-multi-agent`.
- JSON traces, scorecards, run manifests, remediation, Markdown reports, and
  self-contained HTML reports.
- Deterministic validators for artifact integrity and declared benchmark semantics.
- Linux/Python 3.11-3.13 as the primary installed-package contour, with Windows 3.11
  compatibility coverage.
- Reproducible wheel/sdist builds, checksums, GitHub attestations, and exact-subject
  CycloneDX SBOMs for public releases from `v1.0.0` onward.
- A shipped provider-neutral Agent Host V1 contour for canonical,
  authority-free record/replay, deterministic evaluation, explicit Python instrumentation,
  and a validated 48-case no-network quickstart. The CLI does not execute arbitrary hosts
  or tools and does not authenticate its producer or certify an external system; see
  [Agent Host Adapter SDK](docs/agent-host-adapter.md).
- A shipped local Runtime Gateway synthetic contour with closed pre-dispatch policy,
  bounded OpenAI-compatible and stateless MCP endpoints, deterministic synthetic tools,
  privacy-minimized hash-chain audit, dashboard, and hardened source-build Docker Compose.
- Credential-free retained-envelope normalization for OpenAI Responses, Anthropic
  Messages, Google Interactions, and MCP tool calls. These adapters do not make provider
  calls or grant tool-execution authority.

The exact shipped, experimental, planned, and historical surfaces live in
[`docs/current-state.md`](docs/current-state.md) and
[`docs/capability-matrix.md`](docs/capability-matrix.md). Technical v1 gates and honest
non-claims are in [`docs/v1-readiness.md`](docs/v1-readiness.md).

## If you only have one minute

- Run the no-network demo above.
- Read the [committed comparison](examples/comparison-report/README.md).
- Browse the [showcase](docs/showcase/index.md) and
  [scenario matrix](docs/showcase/scenario-matrix.md).
- See [weak spots and findings](docs/showcase/weak-spots-and-findings.md).
- Check [current state](docs/current-state.md) and the
  [project tracker](docs/project-tracker.md).
- Bring your own local or OpenAI-compatible model through
  [Run your model](docs/run-your-model.md).

## Use your own model or runtime

The shortest cross-platform operator path is
[`docs/run-your-model.md`](docs/run-your-model.md). It covers:

1. a no-model deterministic demo;
2. one explicitly authorized OpenAI-compatible model;
3. a deterministic local swarm comparison;
4. a bounded local-model mini-swarm campaign.

Connection details and scenario selection are documented in
[`docs/connect-models.md`](docs/connect-models.md) and
[`docs/test-your-model.md`](docs/test-your-model.md). External runs are prompt-only
self-report checks unless a separately reviewed host/tool adapter exists. A coherent answer
is not evidence of safe tool execution.

## Benchmark and evidence documentation

Start with the curated [documentation map](docs/README.md) if you are not sure which
contract, operator guide, or research page applies to your task.

The README is the front door; deeper contracts live in `docs/`:

| Question | Source of truth |
|---|---|
| What is shipped now? | [Current state](docs/current-state.md) |
| What work is open? | [Project tracker](docs/project-tracker.md) and [roadmap](docs/roadmap.md) |
| Which testing paths are authorized? | [Authorized testing paths](docs/authorized-testing-paths.md) |
| Which system shapes are evaluated? | [Evaluation topologies](docs/evaluation-topologies.md) |
| What boundary model is used? | [Agentic boundary model](docs/agentic-boundary-model.md) |
| How is the corpus expanded? | [Corpus expansion plan](docs/corpus-expansion-plan.md) |
| What do metrics mean? | [Metric contract](docs/metric-contract.md) |
| How are scenarios sequenced? | [Scenario timeline](docs/scenario-timeline.md) |
| How should reports be showcased? | [Showcase checklist](docs/showcase-report-checklist.md) |
| How is evidence promoted? | [Evidence pack format](docs/evidence-pack-format.md) |
| How are changes reviewed? | [Git evidence workflow](docs/git-evidence-workflow.md) |
| How can an external agent host record and evaluate observations? | [Agent Host Adapter SDK](docs/agent-host-adapter.md) |
| How can I run the local policy gateway and MCP/OpenAI-compatible demo? | [Runtime Gateway synthetic contour](docs/runtime-gateway.md) |
| How are provider tool-call envelopes normalized without credentials? | [Provider-neutral tool-call adapters](docs/provider-tool-adapters.md) |
| How can native Ollama output reach Quarantine and a pure Gateway decision without dispatch? | [Native Ollama adapter (published in 1.6.0)](docs/ollama-quarantine-adapter.md) |
| How do optional components exchange validated observations and findings? | [Extension SDK V1](docs/extension-sdk.md) |
| How is an installed extension distribution verified before explicit registration? | [Extension Distribution Discovery V1](docs/extension-distribution-discovery.md) |
| How does an operator approve, list, disable, or plan rollback without automatic code loading? | [Extension Operator Lifecycle V1](docs/extension-operator-lifecycle.md) |
| Which companion contracts already have executable cross-repository adapters? | [Companion Extension adapters](docs/companion-extensions.md) |
| How are weekly public security inputs reviewed without provider lock-in? | [Security Intelligence extension](docs/security-intelligence-extension.md) |
| How can optional packages add patterns without overriding the stable corpus? | [Corpus Pack SDK V1](docs/corpus-pack-sdk.md) |
| How can an operator connect one local model without opening arbitrary tools? | [Controlled local provider/tool-host adapter](docs/controlled-local-adapter.md) |
| How is the reviewed Playbooks Policy Pack evaluated without importing its code? | [Policy Pack V1 extension](docs/policy-pack-extension.md) |

Specialized reviewer paths:

- [Local Prometheus workflow](docs/local-prometheus-workflow.md)
- [Local model profiles](docs/local-model-profiles.md)
- [Multi-agent handoff toy topology](docs/handoff-toy-topology.md)
- [Security audit causal map](docs/security-audit-causal-map-2026-07-15.md)
- [R5 sanitized research status](docs/r5-research-status.md)
- [Project governance](GOVERNANCE.md)

## Standards and portfolio boundaries

The project publishes conservative mappings to OWASP LLM, NIST, and a direct-fit MITRE
ATLAS subset in [`docs/standards-mapping.md`](docs/standards-mapping.md). These are
maintainer-reviewed mappings, not certification or independent standards validation.
Independent review remains public follow-up work.

The related public contract includes:

- [Threat ontology](docs/threat-ontology.md): 26 provider-neutral failure families.
- [Scenario adjudication ledger](docs/scenario-adjudication-ledger.md): 127 bounded source
  units from 13 enumerated builders.
- [Unified event envelope](docs/unified-event-envelope.md): separation of observations,
  data, authority, advisories, decisions, and effects.

The 127 units are **not 127 canonical attacks** and **not a repository-wide total**. The
contract grants no executor, provider, deployment, enforcement, or production authority.

## Public security stack

Agentic Security Harness is the benchmark/evidence layer in a small public defensive
stack:

```text
llm-safety-playbooks -> ai-agent-handoff -> agentic-transfer-verifier -> agentic-security-harness
```

- [`llm-safety-playbooks`](https://github.com/krivonosoff161/llm-safety-playbooks)
  documents practical boundary rules.
- [`ai-agent-handoff`](https://github.com/krivonosoff161/ai-agent-handoff)
  represents task briefs and handoff state as reviewable files.
- [`agentic-transfer-verifier`](https://github.com/krivonosoff161/agentic-transfer-verifier)
  checks provenance, authority, approval, and audit evidence around transfers.
- This repository measures modeled failures and produces validated benchmark artifacts.

The repositories are related but not interchangeable. A playbook is not a runtime control,
a handoff file is not a sandbox, and a passing benchmark is not a production certificate.
The review-only [ecosystem integration candidate](docs/ecosystem-integration-candidate.md)
builds the optional Transfer and Handoff extension wheels and exercises their explicit
approval lifecycle on Ubuntu and Windows; it does not bundle or auto-install them.

## Release and package status

Release [v1.8.0](docs/releases/v1.8.0.md) is published on
[PyPI](https://pypi.org/project/agentic-security-harness/1.8.0/). Its wheel and
sdist match the attested tag build and TestPyPI. The owner-approved production
upload succeeded; Linux 3.12/3.13 and Windows 3.11 passed, but the initial
Linux 3.11 simple-index lookup failed before application checks. The separate
[read-only verification](https://github.com/krivonosoff161/agentic-security-harness/actions/runs/36813028115)
passed all seven cross-index jobs without another upload. This delivers bounded
retained ancestry admission; independent coverage, recovery and the other
foundation questions remain open.

Release [v1.7.1](docs/releases/v1.7.1.md#publication-evidence) is available on
[PyPI](https://pypi.org/project/agentic-security-harness/1.7.1/). Its wheel and
sdist match the attested tag build and TestPyPI. Initial TestPyPI Windows and
PyPI Linux 3.11–3.13 simple-index lookups failed after successful uploads;
the [read-only verification](https://github.com/krivonosoff161/agentic-security-harness/actions/runs/36371342801)
passed all seven subject and cross-platform install jobs without another upload.
This patch does not turn the five open foundation research questions into shipped
features or general security proofs.

Release [v1.7.0](docs/releases/v1.7.0.md#publication-evidence) is available on
[PyPI](https://pypi.org/project/agentic-security-harness/1.7.0/) with wheel and
sdist hashes matching the attested tag build and TestPyPI. Its new garak plan
adapter is explicit and optional. The initial production/read-only workflows
retain Python 3.12/3.13 dependency-lock failures. A separately corrected
[read-only run](https://github.com/krivonosoff161/agentic-security-harness/actions/runs/36306894913)
passed all seven jobs without a second upload. The
[installed-ecosystem example](examples/installed-ecosystem/README.md) provides
version-specific hash-locked routes with positive and negative controls.

The repository also contains the 16-case functional chain, caller-input checks
and bounded local-model evidence above. Those observations retain their exact
1.6.0 installed-subject commitments; publishing 1.7.0 does not rerun or rewrite
them. The example scripts and reports remain repository-owned, with the current
hash-locked install routes documented in their README. Published 1.6.0 subjects
and their embedded documentation remain immutable.

The following prior release evidence is retained without rewriting its outcomes.

Prior release `v1.5.1` is published on
[PyPI](https://pypi.org/project/agentic-security-harness/1.5.1/) and
[GitHub Releases](https://github.com/krivonosoff161/agentic-security-harness/releases/tag/v1.5.1).
[Published-release verification](https://github.com/krivonosoff161/agentic-security-harness/actions/runs/35516400380)
passed exact-subject provenance/index checks and clean Linux/Python 3.11-3.13 plus
Windows/Python 3.11 installation. Publication makes the bounded core and selected passive
distributions installable; it is not automatic activation, production deployment,
enforcement, provider authority, or security certification.

The [v1.5.0 release](docs/releases/v1.5.0.md) adds explicit Quarantine and
advisory-to-Gateway composition plus exact-pinned external Playbooks receipt ingress.
These pure APIs stop at a Gateway decision: they do not dispatch, activate companions,
authenticate producers or infer safe model intent. The [v1.5.1 notes](docs/releases/v1.5.1.md)
retain the initial TestPyPI and Windows PyPI index-lookup failures and successful
read-only recovery; no published package was rebuilt or uploaded again.
Historical v1.5.0 [verification 35507083204](https://github.com/krivonosoff161/agentic-security-harness/actions/runs/35507083204)
and its release artifacts remain unchanged.

- [Release checklist](docs/release-checklist.md)
- [PyPI release process](docs/release-to-pypi.md)
- [Changelog](CHANGELOG.md)
- [Security policy](SECURITY.md)

Independent standards review and a durable second GitHub reviewer are transparent
post-v1 credibility tasks. Their absence does not change deterministic test results, but
independent review is not claimed.

## Development

```bash
python -m pip install -e ".[dev]"
python -m pytest
python -m ruff check .
python -m mypy src tests tools
ash validate examples/
```

Development follows
`idea -> issue -> branch -> implementation -> tests/artifacts -> PR -> GitHub checks -> review gate`.
See [`docs/git-evidence-workflow.md`](docs/git-evidence-workflow.md),
[`CONTRIBUTING.md`](CONTRIBUTING.md), and [`docs/agent-operating-guide.md`](docs/agent-operating-guide.md).

## Responsible use

Use only synthetic, local, owned, or explicitly authorized targets. Do not use this project
for credential theft, persistence, evasion, destructive payloads, or unauthorized systems.
See [`SECURITY.md`](SECURITY.md) and
[`docs/authorized-testing-paths.md`](docs/authorized-testing-paths.md).

## Citation, contributing, and license

- Citation metadata: [`CITATION.cff`](CITATION.cff)
- Governance: [`GOVERNANCE.md`](GOVERNANCE.md)
- Contributing: [`CONTRIBUTING.md`](CONTRIBUTING.md)
- Support: [`SUPPORT.md`](SUPPORT.md)
- License: Apache-2.0, see [`LICENSE`](LICENSE) and [`NOTICE`](NOTICE)
