Metadata-Version: 2.5
Name: polyjev
Version: 0.1.0
Summary: Typed, calibrated decisions from any LLM: yes/no, choice, rating and grounded extraction with honest probabilities, from Claude, GPT, Gemini or open models on your GPUs.
Project-URL: Homepage, https://github.com/PraveenAShukla/polyjev
Project-URL: Documentation, https://praveenashukla.github.io/polyjev
Project-URL: Repository, https://github.com/PraveenAShukla/polyjev
Project-URL: Changelog, https://github.com/PraveenAShukla/polyjev/blob/main/CHANGELOG.md
Author: Praveen Shukla
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: anthropic,calibration,classification,decisions,gemini,jev,llm,logprobs,openai,systemone,vllm
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: httpx>=0.27
Requires-Dist: openai>=1.60
Requires-Dist: pyyaml>=6.0
Requires-Dist: typer>=0.13
Provides-Extra: all
Requires-Dist: anthropic>=0.108; extra == 'all'
Requires-Dist: fastapi>=0.115; extra == 'all'
Requires-Dist: google-genai>=1.22; extra == 'all'
Requires-Dist: python-multipart>=0.0.9; extra == 'all'
Requires-Dist: uvicorn[standard]>=0.30; extra == 'all'
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.108; extra == 'anthropic'
Provides-Extra: gemini
Requires-Dist: google-genai>=1.22; extra == 'gemini'
Provides-Extra: hf
Requires-Dist: accelerate>=1.0; extra == 'hf'
Requires-Dist: torch>=2.3; extra == 'hf'
Requires-Dist: transformers>=4.45; extra == 'hf'
Provides-Extra: server
Requires-Dist: fastapi>=0.115; extra == 'server'
Requires-Dist: python-multipart>=0.0.9; extra == 'server'
Requires-Dist: uvicorn[standard]>=0.30; extra == 'server'
Description-Content-Type: text/markdown

<h1 align="center">polyjev</h1>

<p align="center">
  <b>Typed, calibrated decisions from any LLM.</b><br>
  Ask yes/no, multiple-choice, rating and extraction questions; get back typed answers with honest probabilities,<br>
  from Claude, GPT, Gemini, or open models on your own GPUs.
</p>

<p align="center">
  <a href="https://github.com/PraveenAShukla/polyjev/actions/workflows/ci.yml"><img alt="CI" src="https://github.com/PraveenAShukla/polyjev/actions/workflows/ci.yml/badge.svg"></a>
  <a href="LICENSE"><img alt="Apache-2.0" src="https://img.shields.io/badge/license-Apache--2.0-blue.svg"></a>
  <img alt="Python 3.10+" src="https://img.shields.io/badge/python-3.10%2B-blue.svg">
  <a href="https://colab.research.google.com/github/PraveenAShukla/polyjev/blob/main/examples/quickstart.ipynb"><img alt="Open in Colab" src="https://colab.research.google.com/assets/colab-badge.svg"></a>
</p>

<p align="center">
  <picture>
    <source media="(prefers-color-scheme: dark)" srcset="docs/assets/bench/confidence-vs-accuracy-dark.png">
    <img alt="Across 7 open models, raw confidence sits far above accuracy; after calibration it lands next to it" src="docs/assets/bench/confidence-vs-accuracy.png" width="760">
  </picture>
</p>

LLMs are great judges, classifiers and extractors, but they give you **text**, and
when they sound sure they are often wrong. polyjev turns any model into a
*decision model*:

- **Always valid.** Answers are read as probabilities over *your* options, so
  there is nothing to parse and no answer outside your schema.
- **Honest numbers.** Every answer has a probability distribution and a
  confidence. Option shuffling reduces position bias, and one-command
  calibration brings "90% sure" close to right 90% of the time
  ([benchmark](bench/README.md)).
- **Grounded extraction.** `Span` answers are always substrings of the input,
  with character offsets. The model cannot invent a value.
- **Any model.** Anthropic, OpenAI, Gemini, OpenRouter and every
  OpenAI-compatible server (vLLM, SGLang, Ollama, llama.cpp), or a Hugging Face
  model loaded straight onto your GPU. One string switches between them.
- **A drop-in Jev API.** `polyjev serve` speaks `POST /v1/systemone`, the API of
  TypeSafe's [Jev](https://www.datacamp.com/blog/system-one-models-jev), so Jev
  clients work against your own models.

```python
import polyjev as pj

judge = pj.Polyjev("vllm/local")   # or "anthropic/claude-opus-5", "openai/gpt-5.6-luna", "hf/Qwen/Qwen3-4B-Instruct-2507"
d = judge.decide(
    {"ticket": "Everything is down and we have a demo at noon."},
    {
        "urgent": pj.Noul("Does the customer need a reply within the hour?"),
        "team": pj.Choice("Which team owns this?", {"billing": None, "outage": "service down", "feature": None}),
        "tone": pj.Score("How angry is the customer?", ["calm", "annoyed", "furious"]),
        "order": pj.Span("the order number, if the customer gives one"),
    },
)
d["urgent"].p              # 0.999
d["team"].value            # 'outage'   (pass an Enum class and you get Enum members back)
d["team"].probabilities    # {'billing': 8e-14, 'outage': 1.0, 'feature': 2e-12}
d["tone"].level            # 'annoyed'
d["order"].found           # False: there is no order number, and polyjev will not invent one
```

<sub>Raw output of Qwen3-4B-Instruct-2507 on vLLM (one RTX 5000 Ada): about 60 ms for all four
questions. Instruction-tuned models are this sure of themselves; the
[benchmark](#honest-probabilities-the-benchmark) shows how often they should be.</sub>

## Install

```bash
pip install "polyjev[all]"     # library + server + Claude + Gemini
pip install "polyjev[hf]"      # + in-process Hugging Face models (torch)
```

No key or GPU? `fake/demo` is a deterministic stand-in model, and the
[Colab notebook](https://colab.research.google.com/github/PraveenAShukla/polyjev/blob/main/examples/quickstart.ipynb)
runs a real model on a free GPU.

```bash
polyjev ask fake/demo "Invoice #A-1042, total due \$1,234.56 by 2024-03-15" \
  --noul "overdue=Is the invoice overdue?" --span "id=the invoice id" --spans "dates=every date"
```

## What people use it for

| | example |
|---|---|
| **LLM-as-a-judge** that knows when it is unsure: auto-accept confident grades, send the rest to a human | [`examples/llm_judge.py`](examples/llm_judge.py) |
| **Guardrails**: classify, then redact personal data by exact offsets | [`examples/guardrails.py`](examples/guardrails.py) |
| **Routing** requests to tools or agents, and asking a clarifying question below a confidence threshold | [`examples/router.py`](examples/router.py) |
| **RAG**: is the answer in the context? Extract it with a highlightable citation | [`examples/rag_answerable.py`](examples/rag_answerable.py) |
| **Triage** and tagging at scale, with `decide_many` and bounded concurrency | [`examples/async_batch.py`](examples/async_batch.py) |
| **Vision**: questions about images and webcam frames | [`examples/images.py`](examples/images.py) |

## Questions and answers

| type | you ask | you get |
|---|---|---|
| `Noul` | a yes/no question | `p` = P(yes), `value`, `confidence` |
| `Choice` | one of 2–26 options (strings, a dict with descriptions, or an Enum) | `value`, `probabilities` over every option, `confidence` |
| `Score` | a level on an ordered scale | `level`, `expected` (probability-weighted level), `probabilities` |
| `Span` / `Spans` | one value / every value of a kind to copy from the input | `text`, `start`, `end`, `confidence`, or `found=False` |

Questions can depend on each other (`depends_on`), be asked only when an
earlier answer matches (`ask_if`), and take images. Extra reads (`samples`)
average over shuffled option orders; `think` lets the model reason first. See
[how it works](docs/concepts.md).

## Honest probabilities: the benchmark

[polyjev-bench](bench/README.md) runs 1,777 questions from six public datasets
(BoolQ, RTE, MNLI, AG News, MMLU, SST-5) through 7 open models on one GPU,
calibrating on one half and reporting on the other:

| model | accuracy | mean confidence | ECE raw → calibrated | decisions / s |
|---|---:|---:|---:|---:|
| Gemma-4-12B-it | **0.810** | 0.974 | 0.173 → 0.064 | 30 |
| Qwen3-14B (FP8) | 0.787 | 0.971 | 0.189 → 0.070 | 18 |
| Qwen3-8B | 0.766 | 0.973 | 0.223 → 0.084 | 58 |
| Qwen3-4B-Instruct-2507 | 0.763 | 0.980 | 0.222 → 0.079 | **107** |
| Phi-4-mini-instruct | 0.706 | 0.821 | 0.115 → **0.060** | 85 |
| Granite-3.3-8B-instruct | 0.704 | 0.922 | 0.235 → 0.087 | 42 |
| OLMo-2-13B-Instruct (FP8) | 0.675 | 0.685 | **0.083** → 0.094 | 27 |

What we learned:
- **Models are over-confident**, and one fitted temperature fixes most of it.
- **Reading logprobs beats asking a model for probabilities.** Verbalized
  confidence cost Qwen3-8B 6 points of accuracy and ran ~10× slower.
- **Shuffling the options** cut raw calibration error by a third.
- **polyjev's logprob reading matches the model's exact logits**: the same
  answer on 99.5% of items.

Full [leaderboard](bench/results/LEADERBOARD.md) · [charts](docs/calibration.md) · rerun it with `bench/run.py`.

Calibrate your own model on your own data:

```bash
polyjev calibrate vllm/local --data labelled.jsonl --out calibration.json   # then `calibration: calibration.json` in polyjev.yaml
```

## Any model

| where the model runs | provider ids | how probabilities are read |
|---|---|---|
| your GPU, in-process | `hf` | exact label-token logits |
| your GPU server | `vllm`, `sglang`, `llamacpp`, `ollama`, `lmstudio` | top-20 logprobs at the answer position |
| hosted APIs | `anthropic`, `openai`, `gemini`, `openrouter`, `together`, `groq`, `deepseek`, `fireworks`, `mistral`, `xai`, `openai-compatible` | logprobs where the API has them, otherwise probabilities from structured output ("verbalized") |

`strategy: auto` probes each model once and falls back on its own. Adding a
provider is one small class ([docs](docs/providers.md)). The open-model paths
(vLLM, in-process Hugging Face) have run on real GPUs. The hosted-API adapters
are tested offline: every request is checked against the vendor SDK's real
signature, at both the newest and the oldest supported SDK version. The
benchmark so far covers open models; hosted results are welcome as PRs.

## Serve it

```bash
polyjev serve --model vllm/local --playground            # http://localhost:8011
curl -s localhost:8011/v1/systemone -H 'content-type: application/json' -d @examples/requests/readme_example.json
```

The request and response shapes, extensions and error bodies follow Jev and
[djev](https://github.com/mmastrac/djev), so existing clients only change the
URL. `pj.Remote(url)` is the typed Python client for any such server. The
built-in playground has a JSON editor, image upload and webcam capture, and
probability bars. [Server docs](docs/server.md).

## Run open models on your own GPUs

- **Workstation:** `docker compose --profile gpu up`
  ([`deploy/`](deploy/compose.yaml)) runs vLLM and polyjev together.
- **HPC / SLURM:**
  - [`setup-env.sh`](deploy/slurm/setup-env.sh) installs the vLLM/torch build that matches your driver.
  - [`polyjev.sbatch`](deploy/slurm/polyjev.sbatch) serves a model and prints the SSH tunnel command.
  - Both were tested on a university cluster, including the gotchas: CUDA 12 drivers, GPUs shared with a display, compute nodes without `nvcc`, and home-directory quotas.
- **Inside any GPU job:** `provider: hf`, no server.

[Deployment guide](docs/deploy.md), with a GPU-memory → model table.

## How it compares

| | polyjev | JSON mode / structured-output libraries | Jev (TypeSafe) | [AnyJev](https://github.com/MorrisZJ/AnyJev) |
|---|---|---|---|---|
| answer always valid | yes | yes | yes | yes |
| probabilities for every option | **yes** | no, one label | yes | yes |
| calibration tooling | yes | no | trained in | yes |
| hosted models (Claude, GPT, Gemini) | yes | yes | no, one model | no |
| open models on your GPUs | yes | yes | no | yes |
| grounded extraction with offsets | yes | no | no (Noul/Choice/Score) | no |
| Jev-compatible HTTP API | yes | no | yes | no |
| open source | Apache-2.0 | varies | no | Apache-2.0 |

## FAQ

**Is this Jev?** No. polyjev is independent and not affiliated with TypeSafe;
"Jev" is their trademark. polyjev implements the publicly documented
`/v1/systemone` request and response shapes, so it can stand in for Jev with
models you choose. Jev itself is one trained model with built-in calibration.

**Why not just ask the model for JSON?** You get a label but not how sure the
model is. The numbers a model writes about itself are less reliable than its
token probabilities (see the benchmark), and a bare label cannot tell you when
to escalate to a human.

**Does it work with Claude?** Yes. Claude has no logprobs, so polyjev asks for a
probability per option through structured output ("verbalized"), and prompt
caching keeps the repeated state cheap. The adapter is tested offline against
Anthropic's SDK, but Claude has not been benchmarked here yet. Calibrate it on
a few hundred of your own examples.

**How fast is it?** With an open model on vLLM, one read is a few output tokens:
about 100 ms per decision, and 20–100 decisions per second on one workstation
GPU (see the benchmark). Hosted APIs take their usual round trip per question, asked in parallel.

**What does `samples: "auto"` do?** It reads once, and re-reads with shuffled
options only when the first answer is uncertain. Confident questions cost one
call.

## Project layout

```text
src/polyjev/   library: types, validation, read strategies, engine, spans, client, backends/, server/, CLI
bench/         polyjev-bench: tasks, runner, report, results, SLURM scripts
deploy/        Dockerfile, compose (+ vLLM), SLURM recipes
docs/          documentation site (mkdocs)
examples/      runnable use cases, configs, Colab notebook
tests/         offline tests with every provider SDK faked, plus a real-model test for `hf`
```

## Credits

polyjev follows the API of TypeSafe's **Jev** and the details of Matt Mastracci's
**[djev](https://github.com/mmastrac/djev)** / **[djev-spark](https://github.com/mmastrac/djev-spark)**
(validation rules, `samples: "auto"`, dependencies, span semantics). Its
option-shift debiasing follows **[AnyJev](https://github.com/MorrisZJ/AnyJev)**'s
cyclic shifts.

If polyjev helps your work, a ⭐ helps others find it, and you can [cite it](CITATION.cff).

## Roadmap

- Benchmark results for hosted models (Claude, GPT-5.x, Gemini)
- `pj.Form`: decision schemas declared as classes
- Joint mode: every question in one call, for APIs that bill per request
- A TypeScript client
- Hidden-state probes for in-process models (AnyJev-style)

## Contributing

Bug reports, new providers, benchmark results for more models, and docs fixes are
very welcome. See [CONTRIBUTING.md](CONTRIBUTING.md).

## License

[Apache-2.0](LICENSE)
