Metadata-Version: 2.5
Name: btx_skill_jev_judge
Version: 0.2.4
Summary: Batch-judge items with the TypeSafe Jev API: run, summarize and check-key from the command line
Project-URL: Homepage, https://github.com/bitranox/btx-skill-jev-judge
Project-URL: Repository, https://github.com/bitranox/btx-skill-jev-judge.git
Project-URL: Issues, https://github.com/bitranox/btx-skill-jev-judge/issues
Author-email: bitranox <bitranox@gmail.com>
License: MIT
License-File: LICENSE
Keywords: batch,claude-code-skill,cli,jev,llm-judge,typesafe
Classifier: Environment :: Console
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: click>=8.5.0
Requires-Dist: httpx2>=2.13.1
Requires-Dist: lib-cli-exit-tools>=2.3.4
Requires-Dist: lib-layered-config>=6.1.1
Requires-Dist: lib-log-rich>=6.3.7
Requires-Dist: orjson>=3.12.0
Requires-Dist: pydantic>=2.13.5
Requires-Dist: rich-click>=1.9.9
Provides-Extra: dev
Requires-Dist: bandit>=1.9.4; extra == 'dev'
Requires-Dist: build>=1.6.1; extra == 'dev'
Requires-Dist: click>=8.5.0; extra == 'dev'
Requires-Dist: httpx2>=2.13.1; extra == 'dev'
Requires-Dist: hypothesis>=6.168.3; extra == 'dev'
Requires-Dist: import-linter>=2.15; extra == 'dev'
Requires-Dist: jaraco-context>=6.1.2; extra == 'dev'
Requires-Dist: pip-audit>=2.10.1; extra == 'dev'
Requires-Dist: pynacl>=1.6.2; extra == 'dev'
Requires-Dist: pyright[nodejs]>=1.1.414; extra == 'dev'
Requires-Dist: pytest-cov>=7.1.0; extra == 'dev'
Requires-Dist: pytest>=9.1.1; extra == 'dev'
Requires-Dist: python-multipart>=0.0.32; extra == 'dev'
Requires-Dist: rtoml>=0.13.0; extra == 'dev'
Requires-Dist: ruff>=0.16.10; extra == 'dev'
Requires-Dist: textual>=8.2.8; extra == 'dev'
Requires-Dist: twine>=7.0.0; extra == 'dev'
Requires-Dist: urllib3>=2.8.0; extra == 'dev'
Requires-Dist: virtualenv>=21.14.2; extra == 'dev'
Requires-Dist: wheel>=0.48.0; extra == 'dev'
Description-Content-Type: text/markdown

# btx-skill-jev-judge

<!-- Badges -->
[![CI](https://github.com/bitranox/btx-skill-jev-judge/actions/workflows/default_cicd_public.yml/badge.svg)](https://github.com/bitranox/btx-skill-jev-judge/actions/workflows/default_cicd_public.yml)
[![CodeQL](https://github.com/bitranox/btx-skill-jev-judge/actions/workflows/codeql.yml/badge.svg)](https://github.com/bitranox/btx-skill-jev-judge/actions/workflows/codeql.yml)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Open in Codespaces](https://img.shields.io/badge/Codespaces-Open-blue?logo=github&logoColor=white&style=flat-square)](https://codespaces.new/bitranox/btx-skill-jev-judge?quickstart=1)
[![PyPI](https://img.shields.io/pypi/v/btx-skill-jev-judge.svg)](https://pypi.org/project/btx-skill-jev-judge/)
[![PyPI - Downloads](https://img.shields.io/pypi/dm/btx-skill-jev-judge.svg)](https://pypi.org/project/btx-skill-jev-judge/)
[![Code Style: Ruff](https://img.shields.io/badge/Code%20Style-Ruff-46A3FF?logo=ruff&labelColor=000)](https://docs.astral.sh/ruff/)
[![codecov](https://codecov.io/gh/bitranox/btx-skill-jev-judge/graph/badge.svg)](https://codecov.io/gh/bitranox/btx-skill-jev-judge)
[![security: bandit](https://img.shields.io/badge/security-bandit-yellow.svg)](https://github.com/PyCQA/bandit)

A Claude Code plugin with one skill, `jev-judge`, and the Python package behind it,
`btx-skill-jev-judge`. When the same judgment repeats over many items, such as labeling 600 support
tickets, Claude hands it to [TypeSafe's Jev](https://docs.typesafe.ai) through a tested
command-line tool. That beats reading every item into its own context or calling a large model
once per item.

Jev returns typed answers with probabilities: yes/no, one of a set of options, or a position on a
scale. Input is billed per token and output is free; `run` reports the run's cost, estimated from the
per-million-token price set in the code (`PRICE_PER_MTOK` in `domain/summary.py`).

## Install in Claude Code

The repository is its own plugin marketplace, so installing takes two steps: register the
marketplace once, then install the plugin from it. Set up the [prerequisites](#prerequisites)
(uv and a TypeSafe key) first; the plugin installs without them, but the skill cannot run.

### In a Claude Code session

```text
/plugin marketplace add bitranox/btx-skill-jev-judge
/plugin install btx-skill-jev-judge@btx-skill-jev-judge
```

`/plugin install` opens the plugin panel, where you choose the scope:

| Scope   | Who gets it                   | Recorded in                         |
|---------|-------------------------------|-------------------------------------|
| user    | you, in every project         | `~/.claude/settings.json`           |
| project | everyone working in this repo | `.claude/settings.json` (commit it) |
| local   | you, in this repo only        | `.claude/settings.local.json`       |

With project scope, committing the settings file turns the plugin on for your collaborators but
does not download it: each of them runs the install command once.

Closing the panel loads the plugin into the open session. If it does not appear, run
`/reload-plugins`.

### From your shell

Useful in a setup script:

```bash
claude plugin marketplace add bitranox/btx-skill-jev-judge
claude plugin install btx-skill-jev-judge@btx-skill-jev-judge                  # user scope
claude plugin install btx-skill-jev-judge@btx-skill-jev-judge --scope project  # or project scope
```

Plugins installed this way load in the next session, or after `/reload-plugins` in an open one.

### Check that it works

`claude plugin list` shows `btx-skill-jev-judge@btx-skill-jev-judge`. In a session, describe a matching task
("tag these 600 support tickets by product area") and Claude loads the skill, or invoke it
directly with `/btx-skill-jev-judge:jev-judge`. Its first step runs `check-key`,
which tells you whether the TypeSafe key is usable without printing it.

### Updates

Auto-update is off by default for marketplaces other than Anthropic's own. Update by hand:

```bash
claude plugin marketplace update btx-skill-jev-judge
claude plugin update btx-skill-jev-judge@btx-skill-jev-judge
```

or turn auto-update on in `/plugin`, on the **Marketplaces** tab. An open session keeps the version
it loaded until you run `/reload-plugins`.

### Uninstall

```bash
claude plugin uninstall btx-skill-jev-judge@btx-skill-jev-judge  # add --scope project for a project install
claude plugin marketplace remove btx-skill-jev-judge       # also uninstalls its plugins
```

### From a local clone

To try a change before it is pushed, register the directory instead of the GitHub repository.
Start a relative path with `./`, or Claude Code reads it as `owner/repo`:

```text
/plugin marketplace add ./btx-skill-jev-judge
/plugin install btx-skill-jev-judge@btx-skill-jev-judge
```

## Prerequisites

- [uv](https://docs.astral.sh/uv/). It fetches Python and the CLI's dependencies on first use.
- A TypeSafe API key from https://console.typesafe.ai/keys. Lookup order: an exported
  `TYPESAFE_API_KEY`, then a `.env` in the current directory or a parent, then
  `~/.credentials/typesafe.key` with mode 600 (see [SECURITY.md](SECURITY.md)). To create the file,
  run this, then paste the key in with an editor:

  ```bash
  mkdir -p -m 700 ~/.credentials && install -m 600 /dev/null ~/.credentials/typesafe.key
  ```

## Installation

The command-line tool is the PyPI package `btx-skill-jev-judge`. It installs three command names for
one program: `jev-judge`, `btx-skill-jev-judge` and `btx_skill_jev_judge`.

```bash
uvx --from btx-skill-jev-judge jev-judge --help   # run once, nothing installed
uv tool install btx-skill-jev-judge               # or install it on your PATH
jev-judge --version
```

pip, pipx and other methods are in [INSTALL.md](INSTALL.md). Python 3.10 or newer is required.

## Quick Start

Write one JSON object per item, with an `id` and a `state` of named fields:

```json
{"id": "t-1", "state": {"title": "Login loops after SSO", "body": "I sign in and land on the login page again."}}
{"id": "t-2", "state": {"title": "Add dark mode", "body": "Please add a dark theme."}}
```

Write the questions once. `noul` is Jev's yes/no type and answers with the probability of yes:

```json
[{"id": "is_bug", "type": "noul", "instructions": "Does `body` describe something that is broken?"}]
```

Then check the key, try ten items, run all of them and summarize:

```bash
jev-judge check-key
jev-judge run --items items.jsonl --questions questions.json --out rows.jsonl --pilot 10
jev-judge run --items items.jsonl --questions questions.json --out rows.jsonl
jev-judge summarize --rows rows.jsonl
```

| Command     | Purpose                                                                                     |
|-------------|---------------------------------------------------------------------------------------------|
| `run`       | Asks the same questions about every item and writes one result row per item                 |
| `summarize` | Shows each question's spread, flags questions whose answers never move, lists rows to check |
| `check-key` | Reports whether a usable key is configured and where it came from, without printing the key |

`run`, `summarize` and `check-key` take `--json` for an `{ok, command, data, skipped}` envelope on stdout, or
`--json-bare` for the data alone (also on failure). Diagnostics always go to stderr.

| Exit code | Meaning                                                                                 |
|-----------|-----------------------------------------------------------------------------------------|
| 0         | Yes: every item answered, no flat question, a key is present                            |
| 1         | No: some row failed, a question looks flat, or `check-key` found no key                 |
| 2         | Usage, input or IO error (also an out-of-range flag); stderr says `jev-judge: <reason>` |
| 78        | Broken configuration, such as `attempts` above 10; stderr says `Error: <reason>`        |

A click usage error (a bad option type, an unknown option, `--json` together with `--json-bare`)
exits 2 and prints no JSON, even under `--json-bare`.

Before any item leaves the machine, every string in its state is redacted: common secret formats
such as tokens, private keys and passwords in URLs, plus the API key itself. Requests share one
rate limiter; the default 20 requests per second stays under Jev's documented 40, and `rate`
is configurable. Transient failures (rate
limits, overloads, 5xx errors, timeouts, dropped connections) are retried with backoff, and
`Retry-After` is respected. Any other failure becomes a row that records the reason.
`<command> --help` lists every option.

## Configuration

Defaults for `run` and `summarize` live in the `[judge]` and `[summary]` sections of a layered
configuration (files, `.env`, environment variables). A command-line flag always wins. The API key
(exported `TYPESAFE_API_KEY`, else a `.env` in the current directory or a parent, else
`~/.credentials/typesafe.key`; see [SECURITY.md](SECURITY.md)) is never a configuration value: a
`judge.api_key` entry is refused with exit 78. Environment variables look like
`BTX_SKILL_JEV_JUDGE___JUDGE__RATE=5`. Every key, default and layer is in [CONFIG.md](CONFIG.md).

## Architecture

The package follows a layered ports-and-adapters layout, and `import-linter` enforces it in the
test run:

| Layer       | Package        | Holds                                                           |
|-------------|----------------|-----------------------------------------------------------------|
| Composition | `composition/` | Wires adapters to ports                                         |
| Adapters    | `adapters/`    | CLI, Jev HTTP client, key lookup, files, configuration, logging |
| Application | `application/` | The judge use case and the port protocols                       |
| Domain      | `domain/`      | Questions, items, rows, redaction, summary; no I/O              |

`pyproject.toml` defines two contracts. "Clean Architecture layers" lets each layer import only
the layers below it, in the order composition, adapters, application, domain. "Domain is pure"
also keeps the domain from importing adapters or composition. Tests use real seams: a loopback
HTTP stub stands in for Jev, `CliRunner` drives the CLI, and a fake client drives the use case. See
[docs/systemdesign/module_reference.md](docs/systemdesign/module_reference.md).

## Further Documentation

- [INSTALL.md](INSTALL.md): every way to install the CLI
- [CONFIG.md](CONFIG.md): configuration sections, layers and environment variables
- [DEVELOPMENT.md](DEVELOPMENT.md): the dev loop, gates and releasing
- [CONTRIBUTING.md](CONTRIBUTING.md): how to send a change
- [SECURITY.md](SECURITY.md): key handling, redaction and reporting a vulnerability
- [docs/systemdesign/module_reference.md](docs/systemdesign/module_reference.md): module map
- [docs/measurements.md](docs/measurements.md): one live check against the real API
- [CHANGELOG.md](CHANGELOG.md): what changed in each release
- [ai-stance.md](ai-stance.md) and [ai-transparency.md](ai-transparency.md): why and where AI was used here

## License

MIT
