Metadata-Version: 2.4
Name: neens-eval
Version: 0.3.0
Summary: Neens pre-prod evaluation runner — the CI one-liner (PULL/PUSH) with gate-as-code.
Project-URL: Homepage, https://neens.ai
Project-URL: Documentation, https://app.neens.ai/docs/guides/preprod-evals/
Author: Neens
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: agents,ci,evals,evaluation,gate,llm,neens,opentelemetry,pre-prod,regression-testing
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Quality Assurance
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == 'dev'
Description-Content-Type: text/markdown

# neens-eval — pre-prod evaluation runner (PULL + PUSH)

The thin CI glue that turns a Neens **pre-prod evaluation** into a one-liner. Before a change
ships, replay a golden dataset's frozen prompts against your **candidate** agent, capture the
resulting OpenTelemetry traces, let Neens score them with your existing judges, and **gate** the
build on regressions vs a baseline.

Two models are supported. In the **PULL** model (below) *your* harness runs *your* agent per prompt
and emits traces Neens correlates. In the **PUSH** model ([`neens eval push`](#push-model--neens-calls-your-agent-endpoint-no-harness))
you register your agent as an `agent_http` connection and Neens calls it per prompt itself — no
harness, no OTel wiring.

This runner does **not** try to be a universal agent-caller. In the pull model *your* harness
runs *your* agent; the runner just:

1. fetches the run's frozen prompts,
2. runs **your command once per prompt** with the Neens correlation attributes injected into the
   environment,
3. starts the run, polls until scoring finishes, fetches the gate, and
4. returns an **exit code** (0 = gate passed) so CI can gate on it.

> **Requirement (be honest about it):** your agent must already be OpenTelemetry-instrumented and
> exporting traces to the Neens OTLP receiver. Neens is your observability tool — it correlates the
> traces your agent emits; it does not instrument your agent for you.

Zero runtime dependencies (stdlib only): installing it adds nothing to a CI image beyond
Python 3.10+.

**Documentation:** [Pre-prod evaluations guide](https://app.neens.ai/docs/guides/preprod-evals/) ·
[Framework quickstarts](https://app.neens.ai/docs/guides/quickstarts/) (copy-paste OTel config for
LangGraph, CrewAI, OpenAI Agents SDK, Pydantic AI, Claude Agent SDK, Vercel AI SDK) ·
[neens.ai](https://neens.ai). A TypeScript/Node build with identical flags and exit codes is
published as [`@neens/eval`](https://www.npmjs.com/package/@neens/eval).

## Install

```bash
pip install neens-eval
neens --version
```

This installs the `neens` console script (also runnable as `python -m neens_eval.cli …`).
Pin the version in CI (`pip install "neens-eval==X.Y.Z"`) so a new release never changes a gate
under you.

## The CI one-liner

```bash
neens eval run \
  --base-url "$NEENS_BASE_URL" \
  --api-key  "$NEENS_API_KEY" \
  --run-id   "$RUN_ID" \
  -- python -m my_agent --answer   # <-- YOUR agent command, after `--`
```

Everything after `--` is your command. It is invoked **once per frozen prompt**. The exit code is
the gate verdict, so a GitHub Actions / GitLab step fails automatically when the gate fails:

```yaml
- name: Pre-prod eval gate
  env:
    NEENS_BASE_URL: https://neens.example.com
    NEENS_API_KEY: ${{ secrets.NEENS_PROJECT_KEY }}   # a nk_live_… project key
    # Point YOUR agent's OTel exporter at the Neens receiver:
    OTEL_EXPORTER_OTLP_ENDPOINT: https://neens.example.com
    OTEL_EXPORTER_OTLP_TRACES_ENDPOINT: https://neens.example.com/v1/traces
  run: |
    neens eval run --run-id "$RUN_ID" -- python -m my_agent --answer
```

### Create the run in the same step

If you don't pre-create the run, `--create` makes one against a dataset's golden version:

```bash
neens eval run --create \
  --dataset "$DATASET_ID" --version-label "$GIT_SHA" \
  --baseline prod:7d --max-regressions 0 --min-pass-rate 0.9 \
  -- python -m my_agent --answer
```

## How each prompt reaches your command

For every item the runner sets, and then runs your command:

| Channel | What your agent receives |
| --- | --- |
| **stdin** | the prompt `input` (piped in) |
| `$NEENS_ITEM_INPUT` | the same prompt `input` |
| `$NEENS_EVAL_RUN_ID` | the pre-prod run id |
| `$NEENS_DATASET_ITEM_ID` | the frozen prompt id |
| `$NEENS_VERSION_LABEL` | the candidate version label |
| `$OTEL_RESOURCE_ATTRIBUTES` | the three correlation attrs, **merged** into any existing value |

Read the prompt from stdin **or** `$NEENS_ITEM_INPUT` — whichever suits your harness.

## How the OTel correlation attributes get injected

Neens links an incoming trace to its pre-prod item by three OpenTelemetry attributes:

```
neens.eval_run_id       # which pre-prod run
neens.dataset_item_id   # which frozen prompt
neens.version_label     # the candidate version being replayed
```

The runner injects them **without any change to your agent code** via the standard
`OTEL_RESOURCE_ATTRIBUTES` environment variable (comma-separated `key=value`), which every OTel
SDK merges into the trace resource. If you already set `OTEL_RESOURCE_ATTRIBUTES` (e.g.
`service.name=my-agent,deployment.environment=ci`), the runner **merges** — it keeps your keys and
adds/overrides only the three `neens.*` ones. So every span your agent emits during that
invocation is tagged for correlation automatically.

## Pointing your agent's traces at Neens

Your agent's OTel exporter must send to the Neens OTLP receiver (`POST /v1/traces`). Set the
standard OTel env vars **in your CI environment** (the runner passes the environment through to
your command):

```bash
export OTEL_EXPORTER_OTLP_ENDPOINT=https://neens.example.com
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=https://neens.example.com/v1/traces
# If your OTel SDK sends auth via headers, use your Neens project key:
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer $NEENS_API_KEY"
```

Use whichever endpoint/protocol variables your OTel SDK honors — Neens accepts OTLP/HTTP
(protobuf and JSON) at `/v1/traces`.

## PUSH model — Neens calls your agent endpoint (no harness)

The pull model above needs *your* harness to run *your* agent per prompt. If instead you register
your agent as an **`agent_http`** LLM connection in Neens, Neens can call that endpoint itself —
once per frozen prompt — capture each response, score it, and gate. No local command, no
subprocess, no OTel wiring: you just point the SDK at the run.

```bash
neens eval push \
  --base-url "$NEENS_BASE_URL" \
  --api-key  "$NEENS_API_KEY" \
  --dataset  "$DATASET_ID" \
  --version-label "$GIT_SHA" \
  --agent-connection "$AGENT_CONNECTION_ID" \
  --baseline prod:7d --max-regressions 0 --min-pass-rate 0.9
```

`neens eval push` always **creates** the run (`runner_mode="push"`) against the dataset's golden
version, triggers Neens to call `--agent-connection` per prompt, polls until scoring finishes,
fetches the gate, and returns the **same exit-code contract** (0 = passed, 1 = failed, 2 = error).
Under the inline queue the run is already terminal on trigger, so there is nothing to poll. Use it
as a CI step exactly like `neens eval run`:

```yaml
- name: Pre-prod eval gate (push)
  env:
    NEENS_BASE_URL: https://neens.example.com
    NEENS_API_KEY: ${{ secrets.NEENS_PROJECT_KEY }}   # a nk_live_… project key
  run: |
    neens eval push \
      --dataset "$DATASET_ID" --version-label "$GIT_SHA" \
      --agent-connection "$AGENT_CONNECTION_ID" --max-regressions 0
```

`--no-gate` and `--json` behave as in `run` (the push JSON carries no per-item `items`, since Neens
drives the calls). Flags:

```
neens eval push --dataset ID --version-label LABEL --agent-connection ID [options]

target:
  --dataset ID                dataset to snapshot (its golden version, unless…)
  --dataset-version ID        …an explicit version
  --version-label LABEL       candidate version label (stamped on the run)
  --agent-connection ID       registered agent_http connection Neens calls per prompt (required)
  --name NAME                 run name (defaults to the version label)
  --baseline SPEC             prod | prod:<range> | run:<preprod_run_id>
  --max-regressions N         gate: fail if regressions > N
  --min-pass-rate R           gate: fail if pass rate < R (0..1)

policy (see Gate-as-code policy below):
  --gate-policy FILE          gate-as-code policy file (env NEENS_GATE_POLICY)
  --max-cost-delta-pct PCT    --max-latency-delta-pct PCT
  --min-avg-score R           --max-abs-failures N

connection:
  --base-url URL              Neens API origin         (env NEENS_BASE_URL)
  --api-key KEY               nk_live_… project key    (env NEENS_API_KEY)
  --project-id ID             X-Neens-Project-Id      (env NEENS_PROJECT_ID; usually unneeded)

execution:
  --poll-interval SECONDS     status poll cadence (default 5)
  --poll-timeout SECONDS      overall scoring budget (default 1800)
  --no-gate                   report the gate but exit 0
  --json                      machine-readable result on stdout
```

## Gate-as-code policy

The server gate (`--max-regressions` + `--min-pass-rate`) stays authoritative for its two knobs. On
top of it you can commit a **gate-as-code policy file** that enforces far richer conditions —
cost/latency/step budgets, per-metric score floors, absolute failure caps — evaluated **client-side**
from data Neens already returns (`/comparison`, `/metrics`, run detail). Point the runner at it with
`--gate-policy` (or `$NEENS_GATE_POLICY`):

```bash
neens eval run --run-id "$RUN_ID" --gate-policy neens-gate.json -- python -m my_agent
```

The file is JSON (or YAML, only if `pyyaml` is installed). This example ships in the source
distribution as `neens-gate.example.json` — copy it into your repo as a starting point:

```json
{
  "gate": { "max_regressions": 0, "min_pass_rate": 0.9 },
  "rules": [
    { "cost":    { "max_avg_usd": 0.05, "max_delta_pct": 20 }, "severity": "warn" },
    { "latency": { "max_avg_ms": 3000, "max_delta_pct": 25 }, "severity": "warn" },
    { "steps":   { "max_avg": 8, "max_delta_pct": 50 }, "severity": "block" },
    { "max_abs_failures": 3, "severity": "block" },
    { "max_regressions": 0, "severity": "block" },
    { "min_pass_rate": 0.9, "severity": "block" },
    { "min_avg_score": 0.8, "severity": "block" },
    { "metric": "faithfulness", "min_avg": 0.8, "severity": "block" },
    { "metric": "toxicity_safety", "min_avg": 0.9, "severity": "warn" }
  ],
  "on_missing_baseline": "warn"
}
```

### The `gate` block

The two server-gate knobs. The SDK passes these to the run at **create** time (`--create` / `push`),
so the SERVER stays authoritative for regressions/pass-rate. CLI flags (`--max-regressions`,
`--min-pass-rate`) override the block. On an existing run (`--run-id`) it's already baked in.

### Rules — dimensions and where each signal comes from

Each rule object carries **exactly one** dimension key plus an optional `"severity"` (`"block"`
default, or `"warn"`).

| Dimension | Threshold keys | Backend signal |
| --- | --- | --- |
| `cost` | `max_avg_usd`, `max_delta`, `max_delta_pct` | avg from `comparison.traits.candidate.avg_cost`; delta/pct from `comparison.deltas.cost`. `avg_cost` averages only the sessions Neens could **price**, and is `null` when a side ran entirely on models with no rate — every cost rule is then reported **`skipped` (unpriced)** with the models named, never silently passed and never mistaken for a missing baseline. Set the rate in Settings → Model pricing. |
| `latency` | `max_avg_ms`, `max_delta_ms`, `max_delta_pct` | avg from `traits.candidate.avg_latency`; delta/pct from `deltas.latencyMs` |
| `steps` | `max_avg`, `max_delta`, `max_delta_pct` | avg from `traits.candidate.avg_steps`; delta/pct from `deltas.steps` |
| `max_abs_failures` | int | `run.aggregate.failed` — ABSOLUTE candidate failures |
| `max_regressions` | int | `len(comparison.regressions)` (fallback `run.aggregate.regressions`) |
| `min_pass_rate` | float | `run.aggregate.passRate` (fallback `comparison.candidate.passRate`) |
| `min_avg_score` | float | `run.aggregate.avgScore` |
| `metric` | `{ "metric": "<key>", "min_avg"?, "max_avg"? }` | the `/metrics` rollup's `avgScore` for that `metricKey` |

`%` deltas aren't returned by the backend — the SDK computes `delta / before * 100` client-side,
guarding `before ∈ {null, 0}`.

**"New failures vs baseline" == regressions.** There is no separate backend "new failure" signal, so
the SDK doesn't invent one: use `max_abs_failures` for an absolute failure cap, `max_regressions` for
the baseline-relative one.

**`metric` rules use the NORMALIZED score (0..1, higher = better)** — the `/metrics` `avgScore` is
already normalized so higher is always better, safety metrics included. So a **safety floor** is a
`min_avg` (e.g. `toxicity_safety` ≥ 0.9), never a raw ceiling.

### `on_missing_baseline`

`block | warn | pass` (default `warn`). Governs **baseline-relative** rules (any `*_delta` /
`*_delta_pct` budget, and `max_regressions`) when the baseline cohort is **absent** (the delta's
`before`/`after` are `null`, or no baseline traits). Absolute rules (`max_avg_*`, `max_abs_failures`,
`min_pass_rate`, `min_avg_score`, `metric` floors) always evaluate — they need no baseline.

* `block` → a baseline-relative rule keeps its own severity (a block rule still fails the build).
* `warn` → it's reported as a warning (never blocks).
* `pass` → it's skipped entirely.

### Block vs warn tiers

* `block` (default) — a failing rule **blocks the build** (contributes to CI **exit 1**).
* `warn` — a failing rule is reported but **never blocks** (exit stays 0 unless another block rule or
  the server gate fails).

The policy only ever ADDS a blocking condition on top of the server gate. With no policy, behavior is
identical to before (server-gate only).

### The server gate speaks `cost` too

The policy's top-level **`gate` block** — the one handed to `create_run` at run-create time —
accepts the **same `cost` family, spelled identically** to the `rules[].cost` dimension, so a budget
can live server-side where every consumer of `GET /preprod-evals/{id}/gate` sees it (the UI, another
CI job, a model sweep), not just this CLI process:

```jsonc
{
  "gate": { "max_regressions": 0, "min_pass_rate": 0.9,
            "cost": { "max_avg_usd": 0.004, "max_delta_pct": -50 } },
  "rules": [ { "cost": { "max_avg_usd": 0.05 }, "severity": "warn" } ]
}
```

**Two evaluation sites, one vocabulary.** `rules[].cost` is evaluated **client-side** by this
runner against the `/comparison` payload — it carries a `severity` and honours
`on_missing_baseline`. `gate.cost` is evaluated **server-side** and is stored with the run; it has
no severity, and a breach fails the gate. The same three keys are accepted in both places (an
unknown one inside `gate.cost` is a `GatePolicyError` at parse time), and the SDK **omits** `cost`
from the posted gate entirely when the policy declares none — it never sends `"cost": null`, which
would fingerprint differently from an absent key server-side and void in-flight model sweeps.

The gate response then carries a `costRules` array (`{rule, status, observed, threshold, reason}`),
absent when the gate declares no cost family. `status` is `pass|fail|skipped`, and **only `fail`
blocks**: a run whose sessions all ran on models with **no price** reports `skipped`, naming the
unpriced models — never a pass. Reading a null average as "$0, therefore under budget" is precisely
the bug per-model pricing exists to kill, and an absent baseline (a delta with nothing to compare
against) is reported as its own, differently-worded skip.

### Override flags

Beyond `--gate-policy`, a few flags inject/override **block** rules (a flag replaces any same-dimension
rule from the file):

```
--max-cost-delta-pct PCT     block if candidate avg cost rises more than PCT% vs baseline
--max-latency-delta-pct PCT  block if candidate avg latency rises more than PCT% vs baseline
--min-avg-score R            block if the run's aggregate avg score is below R (0..1)
--max-abs-failures N         block if absolute candidate failures exceed N
```

With neither a policy file nor any of these flags, `policy=None` — today's behavior exactly.

### CI snippet

```yaml
- name: Pre-prod eval gate (gate-as-code)
  env:
    NEENS_BASE_URL: https://neens.example.com
    NEENS_API_KEY: ${{ secrets.NEENS_PROJECT_KEY }}
  run: |
    neens eval run --run-id "$RUN_ID" \
      --gate-policy neens-gate.json \
      -- python -m my_agent --answer
```

`--json` additionally emits a `policyEval` object (per-rule `dimension`/`severity`/`status`/`observed`/
`threshold`/`reason`, plus `blocked`/`warned`) for a CI step to parse.

## Multi-model sweeps — `neens sweep`

A **sweep** replays ONE frozen golden dataset version against N candidate model arms, k times
each, under identical conditions (same items, same judges, same k, same gate) and reports
**pass^k per arm**. It answers *which model should we ship*, so the CI story is **estimate, then
decide**:

```bash
# 1. What would this cost? Writes nothing, spends nothing.
neens sweep preview \
  --dataset-id ds_checkout \
  --arm "sonnet=conn_sonnet" --arm "haiku=conn_haiku" --arm "local=conn_ollama" \
  --pass-k 3

# 2. Launch it, and wait for the per-arm verdict.
neens sweep start \
  --name "nightly model bake-off" --version-label "$GIT_SHA" \
  --dataset-id ds_checkout \
  --arm "sonnet=conn_sonnet" --arm "haiku=conn_haiku" \
  --pass-k 3 --budget-usd 10 --wait --json

# 3. Or poll it later.
neens sweep get msw-4f2c91ab0d3e

# 4. Get the ANSWER: which arm should we ship, and does it differ per agent?
neens sweep decide --sweep-id msw-4f2c91ab0d3e --bar 0.9
```

Each `--arm` is `LABEL=CONNECTION_ID`, where the connection is a registered `agent_http` LLM
connection whose endpoint runs that model. Two arms on the same connection is rejected — that is
the same model twice, not a comparison. Give one of `--dataset-id` (its frozen golden version),
`--dataset-version-id`, or `--scenario-suite-id`.

### `neens sweep decide` — the verdict

`decide` reads `GET /model-sweeps/{id}/comparison` and prints the **server's** verdict: the winning
arm and why, a per-arm table (pass rate **with its sample size and confidence interval**, pass^k,
cost per case, regressions), and one verdict **per agent** — because "Haiku is good enough" is
routinely true for one agent and false for another.

```
Sweep msw-4f2c91ab0d3e — status=completed
  VERDICT: haiku clears the 90% bar at $0.000600 per case — 81% cheaper than sonnet.
  bar 90.0% (source: gate) | baseline: sonnet (incumbent) | 40 items x k=3
  ARM     MODEL              PASS RATE                    PASS^K    COST/CASE  REGRESSIONS
  sonnet  claude-sonnet-4-6  95.0% (n=40, CI 83.0-99.0%)  3/3 PASS  $0.003100  0
  haiku   claude-haiku-4-5   92.5% (n=40, CI 80.0-98.0%)  3/3 PASS  $0.000600  3
  local   gpt-oss:20b        no data                      0/3 —     unpriced   —
```

Flags: `--bar 0.9` (the pass rate an arm must clear — omit it and the sweep's pinned
`gate.min_pass_rate` is used, else the deployment default; the source is always reported),
`--baseline-arm ID` (the incumbent the savings are measured against — defaults to the arm at
position 0, never the best performer, because a baseline chosen after seeing the results is not a
baseline), `--require-winner`, `--wait`, `--json`.

Two properties this command will not violate:

* **The ranking is the server's.** `decide` renders `verdict`/`agents[].verdict` verbatim — it never
  computes a second ranking, because two rankings that can disagree is worse than one.
* **An unpriced arm never wins "cheapest."** No rate means `unpriced`, never `$0.00`: it is still
  shown, still judged against the bar, and **excluded from the cost ranking with a named reason**.

### Sweep exit codes

| Exit code | Situation |
| --- | --- |
| `0` | The sweep completed — **including** when an arm missed pass^k. A sweep is a model-*selection* decision, not a gate. |
| `0` | Launched without `--wait`: the verdict is not in yet. |
| `1` | `--require-all-arms` was passed and some arm did not hit pass^k (or the sweep has no verdict yet). |
| `2` | The sweep is **`void`** — the arms did not run under identical conditions, so the comparison is refused and `voidReason` is printed. A result we cannot trust must never look like a pass. |
| `2` | The sweep `failed`, or a transport/API error (including a budget refusal, which is a 422). |

For `sweep decide` the same table holds, with `--require-winner` in place of `--require-all-arms`:
a comparison that names **no** clearing arm still exits `0` (naming a winner is a decision, not a
gate) unless `--require-winner` is passed, and a `void` / `comparable: false` comparison is exit
`2` **before any verdict is read** — a void sweep still has per-arm numbers, and reading them first
is exactly how an untrustworthy comparison gets to look like a pass.

`--require-all-arms` (and `--require-winner`) implies `--wait`: you cannot assert every arm hit
pass^k — or require a winner — without having seen every arm finish.

An estimate whose arms include a model with no price in your tenant's table is printed as
**partial**, naming the unpriced models — the total covers the priced arms only and is a floor,
never the cost. Set a rate under Settings → Model pricing for a complete figure.

## Gate exit-code contract

| Exit code | Meaning |
| --- | --- |
| `0` | Gate passed (or `--no-gate` was set) |
| `1` | Gate failed (regressions/pass-rate breached the configured gate) |
| `2` | Run error (no items, backend error, poll timeout, bad arguments) |

`--no-gate` still fetches and prints the gate but always exits `0` (report without blocking CI).
`--json` emits a machine-readable result to **stdout** (human logs go to stderr) for a CI step to
parse.

## Flags

```
neens eval run [target] [policy] [connection] [execution] -- <your agent command>
neens --version            print the installed version

target:
  --run-id ID                 drive an existing run
  --create                    create a run first (needs --dataset + --version-label)
    --dataset ID              dataset to snapshot (its golden version, unless…)
    --dataset-version-id ID   …an explicit version
    --version-label LABEL     candidate version label (also tags traces)
    --name NAME               run name (defaults to the version label)
    --baseline SPEC           prod | prod:<range> | run:<preprod_run_id>
    --max-regressions N       gate: fail if regressions > N
    --min-pass-rate R         gate: fail if pass rate < R (0..1)

policy (see Gate-as-code policy above):
  --gate-policy FILE          gate-as-code policy file (env NEENS_GATE_POLICY)
  --max-cost-delta-pct PCT    --max-latency-delta-pct PCT
  --min-avg-score R           --max-abs-failures N

connection:
  --base-url URL              Neens API origin         (env NEENS_BASE_URL)
  --api-key KEY               nk_live_… project key    (env NEENS_API_KEY)
  --project-id ID             X-Neens-Project-Id      (env NEENS_PROJECT_ID; usually unneeded)

execution:
  --timeout SECONDS           per-item command timeout
  --concurrency N             parallel item invocations (default 1)
  --poll-interval SECONDS     status poll cadence (default 5)
  --poll-timeout SECONDS      overall scoring budget (default 1800)
  --no-gate                   report the gate but exit 0
  --json                      machine-readable result on stdout
```

## Pre-rename environment variables are refused, not ignored

<!-- >>> BEGIN LEGACY-IDENTIFIER REGION (issue #462) <<< — this section's job is to print the dead
     names back to a CI operator, who matches them against a red job log. -->

Neens was previously called Beacon, and the rename to `NEENS_*` was a hard break with no
back-compat aliases. If a variable is still spelled `BEACON_*`, this CLI **refuses to run** (exit
`2`), naming every offender and its replacement:

```console
$ BEACON_GATE_POLICY=neens-gate.json neens eval run --run-id ppr_1 -- ./my-agent
neens: error: legacy BEACON_* configuration detected — this SDK does not read these.

The Beacon → Neens rename was a HARD BREAK with no back-compat aliases. A leftover
BEACON_* name is NOT reported at read time: it is silently ignored and this command
would run on the DEFAULT instead. For BEACON_GATE_POLICY the default is 'no client-side
gate at all', so the run would exit 0 on a candidate that breaches your policy — a green
CI check for a release you meant to block. Refusing to run is the only way you find out.

Rename each of the following, then re-run:

  BEACON_GATE_POLICY  →  NEENS_GATE_POLICY
                         read by `neens eval run/push` as the default for --gate-policy. …
```

The fix is the mechanical one — `BEACON_X` → `NEENS_X`:

```diff
  - name: Pre-prod eval gate
    env:
-     BEACON_BASE_URL: ${{ secrets.NEENS_BASE_URL }}
-     BEACON_API_KEY: ${{ secrets.NEENS_API_KEY }}
-     BEACON_GATE_POLICY: neens-gate.json
+     NEENS_BASE_URL: ${{ secrets.NEENS_BASE_URL }}
+     NEENS_API_KEY: ${{ secrets.NEENS_API_KEY }}
+     NEENS_GATE_POLICY: neens-gate.json
    run: neens eval run --create --dataset ds_1 --version-label "$GITHUB_SHA" -- ./my-agent
```

**Why refuse rather than ignore.** `--gate-policy` defaults to the environment, and an *absent*
value legitimately means "server gate only" — so a stale `BEACON_GATE_POLICY` was not an error, it
simply built no client-side gate, and the run exited `0` on a candidate that breached the policy.
The other variables fail loudly on their own (a missing base URL or API key is a required-argument
error), which is why this one needed a guard and they did not.

The guard covers only the variables this SDK actually uses — `NEENS_GATE_POLICY`, `NEENS_BASE_URL`,
`NEENS_API_KEY`, `NEENS_PROJECT_ID`, and the four it exports into your agent's environment
(`NEENS_EVAL_RUN_ID`, `NEENS_DATASET_ITEM_ID`, `NEENS_VERSION_LABEL`, `NEENS_ITEM_INPUT`) — so an
unrelated `BEACON_*` variable belonging to some other tool never blocks your CI. It runs at CLI
entry for every subcommand; importing `neens_eval` as a library never raises.

If one of those names really does belong to another tool and cannot be unset, list it explicitly:

```bash
export NEENS_IGNORE_LEGACY_ENV="BEACON_API_KEY"   # comma-separated; names only
```

It is deliberately a **list of names, not an on/off switch** (a blanket toggle gets reached for the
moment the guard fires for the right reason), and every ignored name is still reported on stderr —
loudly when it shadows a real Neens variable whose replacement is not set, because that setting is
then running on its default.

<!-- >>> END LEGACY-IDENTIFIER REGION <<< -->

## Library use

PULL model — drive your own agent over an existing run's prompts:

```python
from neens_eval import NeensEvalClient, GatePolicy, run_eval

client = NeensEvalClient("https://neens.example.com", api_key="nk_live_...")
policy = GatePolicy.from_file("neens-gate.json")   # or GatePolicy.from_dict({...})
outcome = run_eval(
    "ppr_123",
    ["python", "-m", "my_agent", "--answer"],
    client=client,
    per_item_timeout=120,
    concurrency=4,
    policy=policy,                                    # client-side gate-as-code (optional)
)
if outcome.policy_eval:
    print("blocked:", outcome.policy_eval.blocked, outcome.policy_eval.reasons())
raise SystemExit(outcome.exit_code)
```

PUSH model — create + invoke a run against a registered agent endpoint (Neens does the calling):

```python
from neens_eval import NeensEvalClient, run_push_eval

client = NeensEvalClient("https://neens.example.com", api_key="nk_live_...")
outcome = run_push_eval(
    client=client,
    dataset_id="ds_1",
    version_label="v2",
    agent_connection_id="conn_abc",          # a registered agent_http connection
    baseline={"kind": "prod_window", "range": "7d"},
    gate={"max_regressions": 0, "min_pass_rate": 0.9},
)
raise SystemExit(outcome.exit_code)
```

## License

Apache-2.0 — see the `LICENSE` file shipped with this package.
