Metadata-Version: 2.4
Name: multivon-eval
Version: 0.20.0
Summary: AI evaluation for teams that ship models to production
Author-email: Multivon <hello@multivon.ai>
License-Expression: Apache-2.0
Project-URL: Homepage, https://multivon.ai
Project-URL: Repository, https://github.com/multivon-ai/multivon-eval
Project-URL: Documentation, https://docs.multivon.ai
Project-URL: Bug Tracker, https://github.com/multivon-ai/multivon-eval/issues
Keywords: llm,evaluation,ai,evals,rag,agents,testing,llm-judge,agent-evaluation
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: anthropic>=0.40.0
Requires-Dist: openai>=1.50.0
Requires-Dist: jsonschema>=4.20.0
Requires-Dist: python-dotenv>=1.0.0
Requires-Dist: rich>=13.0.0
Provides-Extra: litellm
Requires-Dist: litellm>=1.0.0; extra == "litellm"
Provides-Extra: pricing
Requires-Dist: litellm<2,>=1.101.0; extra == "pricing"
Provides-Extra: bertscore
Requires-Dist: bert-score>=0.3.0; extra == "bertscore"
Provides-Extra: requests
Requires-Dist: requests>=2.28.0; extra == "requests"
Provides-Extra: browser
Requires-Dist: playwright>=1.40.0; extra == "browser"
Requires-Dist: requests>=2.28.0; extra == "browser"
Provides-Extra: pytest
Requires-Dist: pytest>=7.0.0; extra == "pytest"
Provides-Extra: google
Requires-Dist: google-genai>=1.0.0; extra == "google"
Provides-Extra: datasets
Requires-Dist: datasets<6,>=5.0.1; extra == "datasets"
Provides-Extra: inspect
Requires-Dist: inspect-ai<0.4,>=0.3.263; extra == "inspect"
Provides-Extra: review
Requires-Dist: scikit-learn<2,>=1.5; extra == "review"
Requires-Dist: scipy<2,>=1.11; extra == "review"
Provides-Extra: otel
Requires-Dist: opentelemetry-sdk<2,>=1.44; extra == "otel"
Requires-Dist: opentelemetry-exporter-otlp-proto-http<2,>=1.44; extra == "otel"
Provides-Extra: gymnasium
Requires-Dist: gymnasium<2,>=1.2; extra == "gymnasium"
Provides-Extra: media
Requires-Dist: Pillow<13,>=12; extra == "media"
Requires-Dist: pypdfium2<6,>=5.13; extra == "media"
Requires-Dist: av<18,>=16; extra == "media"
Provides-Extra: langgraph
Requires-Dist: langgraph>=0.2.0; extra == "langgraph"
Requires-Dist: langchain-core>=0.3.0; extra == "langgraph"
Provides-Extra: openai-agents
Requires-Dist: openai-agents>=0.0.1; extra == "openai-agents"
Provides-Extra: all
Requires-Dist: litellm<2,>=1.101.0; extra == "all"
Requires-Dist: bert-score>=0.3.0; extra == "all"
Requires-Dist: requests>=2.28.0; extra == "all"
Requires-Dist: playwright>=1.40.0; extra == "all"
Requires-Dist: pytest>=7.0.0; extra == "all"
Requires-Dist: google-genai>=1.0.0; extra == "all"
Requires-Dist: langgraph>=0.2.0; extra == "all"
Requires-Dist: langchain-core>=0.3.0; extra == "all"
Requires-Dist: openai-agents>=0.0.1; extra == "all"
Requires-Dist: datasets<6,>=5.0.1; extra == "all"
Requires-Dist: inspect-ai<0.4,>=0.3.263; extra == "all"
Requires-Dist: scikit-learn<2,>=1.5; extra == "all"
Requires-Dist: scipy<2,>=1.11; extra == "all"
Requires-Dist: opentelemetry-sdk<2,>=1.44; extra == "all"
Requires-Dist: opentelemetry-exporter-otlp-proto-http<2,>=1.44; extra == "all"
Requires-Dist: gymnasium<2,>=1.2; extra == "all"
Requires-Dist: Pillow<13,>=12; extra == "all"
Requires-Dist: pypdfium2<6,>=5.13; extra == "all"
Requires-Dist: av<18,>=16; extra == "all"
Provides-Extra: dev
Requires-Dist: pytest>=7.0.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23.0; extra == "dev"
Requires-Dist: ruff>=0.3.0; extra == "dev"
Requires-Dist: build>=1.0.0; extra == "dev"
Requires-Dist: twine>=5.0.0; extra == "dev"
Dynamic: license-file

# multivon-eval

[![PyPI](https://img.shields.io/pypi/v/multivon-eval.svg)](https://pypi.org/project/multivon-eval)
[![Python](https://img.shields.io/badge/python-3.10+-blue.svg)](https://pypi.org/project/multivon-eval)
[![License](https://img.shields.io/badge/license-Apache%202.0-blue.svg)](LICENSE)
[![Tests](https://github.com/multivon-ai/multivon-eval/actions/workflows/test.yml/badge.svg)](https://github.com/multivon-ai/multivon-eval/actions/workflows/test.yml)

**Did your AI application get better—or did the measurement change?**

multivon-eval is a Python library for testing LLM applications, RAG systems,
and agents. Define task-specific cases, grade outputs, compare changes, and
inspect quality failures separately from errors and missing evidence.
Runs and reports stay local; LLM graders call the judge provider you configure.
No hosted account is required.

[Documentation](https://docs.multivon.ai/) · [Examples](examples/README.md) ·
[Benchmarks](benchmarks/README.md) · [Changelog](CHANGELOG.md)

**Current release: 0.20.0 — September 21, 2026.** Python 3.10+, Apache 2.0.
[Migration notes](docs/guides/migration-0-19.mdx).

[Case manifests and trial evidence](docs/guides/versioned-evidence.mdx)
support safer comparisons and regrading. Use Hugging Face for dataset operations,
[Inspect for execution](docs/guides/inspect-integration.mdx), and Multivon's
[acceptance policies](docs/guides/acceptance-policies.mdx) for required checks and
task slices. These features ship in 0.19.0. The
[implementation program](plans/industrial-evaluation.md) tracks unfinished work;
[reuse decisions](plans/reuse-decisions.md) keep the integration boundaries explicit.

## Start in 30 seconds

```bash
pip install multivon-eval
```

Run this complete example. It needs no API key:

```python
from multivon_eval import EvalCase, EvalSuite, ExactMatch

suite = EvalSuite("capital lookup")
suite.add_cases([
    EvalCase(input="Capital of France?", expected_output="Paris"),
    EvalCase(input="Capital of Japan?", expected_output="Tokyo"),
])
suite.add_evaluators(ExactMatch())

# Replace this fixture with your application: a function from str to str.
answers = {"Capital of France?": "Paris", "Capital of Japan?": "Tokyo"}
report = suite.run(
    answers.__getitem__,
    fail_threshold=1.0,
    save_json="results.json",
    save_html="results.html",
    verbose=False,
)
print(f"{report.passed}/{report.evaluated} passed; {report.errors} errors")
# 2/2 passed; 0 errors
```

Open `results.html` to inspect each verdict. The example verifies a small
lookup fixture; it does not establish the quality of a real AI application.
Start your own suite by [defining task success](docs/guides/task-success.mdx).

The [Label Studio review bridge](docs/guides/review-labels.mdx)
exchanges saved text/trace trials and preserves review disagreements and missing
coverage. [Saved-score calibration](docs/guides/review-calibration.mdx) reuses
scikit-learn and SciPy for development fitting and held-out source analysis.
[OpenTelemetry interoperability](docs/guides/otel-evidence.mdx) grades retained
OTLP traces and emits standard evaluation events through your existing SDK.
[Environment outcome checks](docs/guides/environment-outcomes.mdx) reuse Gymnasium
and independently observed state to catch missing writes, duplicate writes and
forbidden changes. The [failure investigation workflow](docs/guides/failure-investigation.mdx)
connects saved trial comparison to Label Studio review and development regression
cases. The [vision grader audit](docs/evaluators/multimodal.mdx) corrects empty-claim
perfect scores, invalid-judgment handling and Anthropic SDK 1.x compatibility.
[Content-bound media](docs/guides/media-evidence.mdx) connects verified image,
PDF, audio and video inputs to native Inspect logs and W3C verdict references.
[Controlled robustness](docs/guides/controlled-robustness.mdx) validates candidate
oracles and preserves rejected or unknown transformations. Hardness filtering
keeps missing measurements separate from model failures.
[Experimental world-model evaluation](docs/guides/world-models.mdx) binds numeric
state forecasts to native environment transitions and tests action effects,
uncertainty and actual planning outcomes. Its tested scope is fully observed
state-space models, not video generation or real robots.
These experimental interfaces ship in 0.19.0 with their documented scope limits.

## A real workflow example

The [document-to-ledger study](benchmarks/industrial/DOCUMENT_RESULTS.md) reuses
CORD receipts, pdfhell generators and Inspect execution. Models write to SQLite;
independent checks verify the saved state. Both models fail the frozen policy
on 39 held-out source documents. The report separates wrong amounts, missing
posts, transcription-only mismatches and integration errors, with raw-log hashes,
cost accounting and a reproducible protocol. It is a sandbox study, not proof
of production readiness or a new dataset.

The [execution-control study](benchmarks/industrial/EXECUTION_CONTROLS_VALIDATION.md)
shows why a successful text reply can coexist with an unfinished task: native
Inspect limits stop six local writes, and independent SQLite checks catch the
missing commits. A completed control verifies the positive path.

The [full TAT-QA financial study](benchmarks/industrial/results/tatqa-financial-2026-09-17/README.md)
tests 1,663 released test questions with the official scorer. Requiring Claude
Sonnet 5 to emit answer and evidence locations together reduced exact match from
74.74% to 72.22% and F1 from 82.76% to 79.96%; both paired context-bootstrap
intervals exclude zero. The joint record also cost 41.70% more. The result argues
for testing answer generation and evidence binding as separate stages, with
missing support remaining indeterminate. It is not a SoTA or customer-validity
claim.

## Why use it

| Need | Available today |
|---|---|
| Catch regressions | Deterministic checks, LLM graders, and saved baseline/proposal comparisons. |
| Detect missing evidence | Separate quality failures, model/judge errors, and skipped checks. Active quality gates block errors and skipped coverage by default. |
| Check the graders | Validate reference answers, measure agreement with reviewed labels, and inspect grader reasons. |
| Measure variability | Repeated trials, flakiness, pass@k/pass^k, and confidence intervals with documented assumptions. |
| Keep results portable | Local JSON, HTML, CSV, JUnit, and optional audit artifacts. |

## Pick your path

| Task | Starting point |
|---|---|
| First offline suite | `multivon-eval init -t quickstart -d my-eval` |
| RAG / question answering | `multivon-eval init -t rag` |
| Agent tool use | `multivon-eval init -t agent` |
| LangGraph agent | `multivon-eval init -t agent-langgraph` |
| OpenAI Agents SDK agent | `multivon-eval init -t agent-openai-sdk` |
| Multi-turn conversations | `multivon-eval init -t conversation` |
| Existing logs | [Score recorded outputs](https://docs.multivon.ai/guides/score-logged-outputs) |
| Unsure what to measure | [Bootstrap a starter suite](https://docs.multivon.ai/guides/bootstrap) |

Bootstrap suggests evaluators and synthetic cases. Its p25 threshold suggestions
are provisional score summaries, not calibration against human acceptance labels.
Review them before using them as release criteria.

## Add an LLM judge

Use deterministic checks for exact requirements and judges for qualities that
need interpretation. Install the relevant provider SDK and set its API key;
configure a local judge if you want to avoid hosted calls.

```python
from multivon_eval import EvalCase, EvalSuite, Faithfulness, JudgeConfig, configure

configure(JudgeConfig(provider="anthropic", model="claude-haiku-4-5"))
suite = EvalSuite("policy answers")
suite.add_cases([EvalCase(
    input="What is the refund window?",
    context="Refunds are available within 30 days of purchase.",
)])
suite.add_evaluators(Faithfulness())

# your_app(prompt) must call your application, including its retrieval step.
report = suite.run(your_app, fail_threshold=0.90, save_json="policy-results.json")
```

Judge providers include Anthropic, OpenAI, Google, Ollama, and LiteLLM;
OpenAI-compatible endpoints support local servers. See
[judge configuration](https://docs.multivon.ai/evaluators/llm-judge) for setup,
historical threshold packs, and the current temperature-forwarding limitation.
A judge score is an estimate to validate on your task, not ground truth.

## Evaluators — 44 across 7 tiers

| Family | Examples | Judge calls? |
|---|---|---|
| Deterministic | `ExactMatch`, `Contains`, `JSONSchemaEval`, `BLEU`, `ROUGE`, `Latency` | No |
| LLM judge | `Faithfulness`, `Hallucination`, `Relevance`, `AnswerAccuracy`, `GEval` | Yes |
| Agent trace | `ToolCallAccuracy`, `ToolArgumentAccuracy`, `TaskCompletion` | Some |
| Conversation | `KnowledgeRetention`, `ConversationCompleteness`, `TurnConsistency` | Yes |
| Compliance checks | `PIIEvaluator`, `SchemaEvaluator` | No |
| Multimodal | `VQAFaithfulness`, `DocumentGrounding` (experimental) | Yes |
| Consistency | `SelfConsistency` | Yes |

[Browse the evaluator reference](https://docs.multivon.ai/evaluators/deterministic).
For tool expectations, `None` means unspecified, `[]` means no calls expected,
and `require_order=True` checks an **ordered subsequence**. A matching trace
alone does not prove the task changed external state correctly.

## Use it in CI

`fail_threshold` checks absolute quality. In 0.17.0, an active gate returns:

| Exit | Meaning |
|---|---|
| `0` | The configured gate passed. |
| `1` | Completed measurements failed the quality threshold. |
| `2` | Evidence is indeterminate: errors, empty runs, or skipped coverage. |

Set `max_error_rate=` explicitly to allow an error budget. Pass `save_json=`,
`save_html=`, or `save_junit_xml=` into `suite.run()` so reports are written
**before** a failing gate raises.

For a saved baseline/proposal comparison:

```bash
multivon-eval compare baseline.json proposal.json --fail-on-regression
```

This also blocks incomplete or unmatched comparisons. It flags case regressions
without waiting for statistical significance; no detected regression is not
proof of equivalence. See [CI/CD](https://docs.multivon.ai/guides/ci-cd) and
[statistical assumptions](https://docs.multivon.ai/guides/statistical-rigor).

## Evidence and limitations

The repository publishes [benchmark scripts and historical results](benchmarks/README.md).
The [first full RAGChecker judge study](benchmarks/industrial/results/ragchecker-judges-2026-09-17/README.md)
measures current `AnswerAccuracy` at Pearson **0.499** against overall human
preference on 280 cases (95% case-bootstrap interval **0.413–0.572**).
It does not establish an advantage over a same-model direct judge: the paired
interval includes zero, and the ranking reverses under a post-hoc formatting
check. QAG used four calls per response versus one for the direct judge.

The [TAT-QA evidence-record study](benchmarks/industrial/results/tatqa-financial-2026-09-17/README.md)
is a separate answer-workflow measurement. Its negative result shows that
structured evidence requirements can reduce answer quality even when location
syntax is usually valid. Public labels, a mutable model alias and annotation
location ambiguity prevent a leaderboard or generalization claim.

One cross-task measurement reported F1 **0.830** on 60 HaluEval summarization
outputs from 30 source examples, using a QA-selected threshold. It is a small,
maintainer-run result with generated hallucination labels; it does not establish
state-of-the-art accuracy. Historical live benchmarks have not been rerun after
the 0.17.0 grader changes.

Confidence intervals do not correct biased labels, correlated cases, or grader
mistakes. Validate your success criteria and inspect errors, skips, and important
task slices. Optional compliance reports organize evidence; they do not certify
legal compliance.

A case with one measured check and another skipped check still counts as evaluated.
The built-in gate blocks wholly skipped cases; enforce per-check coverage separately
when every evaluator is required.

## Useful commands

```bash
multivon-eval validate eval.py                 # check reference outputs against graders
multivon-eval view --dir runs/                 # browse saved reports
multivon-eval bootstrap --product PRODUCT.md --traces TRACES.jsonl
multivon-eval generate --from docs/faq.md --n 20
multivon-eval assess traces.jsonl              # inspect input quality locally
multivon-eval staleness .                      # inspect prompt/case drift
multivon-eval doctor --no-ping --json          # check configuration offline
```

`doctor` exits 0 when clean, 2 when it finds warnings, and 1 when it finds an error.
Use `multivon-eval --help` for all commands.

## Current release — 0.20.0

- Grade content-bound media in the vision evaluators: bytes are verified against the case's descriptor before a judge sees them, and a missing resolver or changed bytes fail instead of scoring.
- Make open-weights reasoning judges usable. Their verdicts were truncated into unparseable working, because the reasoning token floor recognised only OpenAI names and was too low for claim extraction.
- Publish the first calibration rows for a judge that is neither Anthropic's nor OpenAI's, and the first open-weights external-judge measurement.
- Share one provider dispatch between `vision.py` and the multimodal evaluators, which gains them local VLMs without losing native provider evidence.
- Cover the self-hosted judge routes — ollama, an OpenAI-compatible `base_url`, and litellm — with tests.

See the [changelog](CHANGELOG.md) for the measured failures behind each of these.
Earlier [migration notes](docs/guides/migration-0-19.mdx), the [worked document study](benchmarks/industrial/DOCUMENT_RESULTS.md),
and the [changelog](CHANGELOG.md). This release does not establish complete
provider capture, opaque callback compatibility, customer validation or SoTA accuracy.

[Strict agent judgment evidence](docs/evaluators/agent.mdx) and
[declared grader dependencies and native provider evidence](docs/guides/versioned-evidence.mdx)
ship in 0.19.0.
Instrumented SDK calls retain request attempts and complete native usage fields;
an optional SQLite journal preserves dispatched attempts after process death.
Unobserved transports, streaming usage and missing responses stay explicit gaps.
The [accounting workflow](docs/guides/provider-accounting.mdx) reuses
LiteLLM pricing and rejects incomplete provider budget evidence. Recorded judge
subtotals do not establish a complete run cost.
Target/runner snapshots and `declare_target` expose intended interventions and
observed changes during execution; regrading preserves original target evidence.
[Execution controls](docs/guides/inspect-integration.mdx#execution-limits-and-completion-evidence-development)
reuse native Inspect limits and retain stop reasons. Saved-output regrading cannot
erase an undeclared stop or invalidation; bounded-task acceptance requires an
explicit policy and outcome checks. Native async cancellation drains owned tasks,
with synchronous worker-thread limits documented.
[Declared Inspect tasks](docs/guides/inspect-integration.mdx#retry-compatibility-development)
bind retry compatibility to task inputs, grader settings and dependency revisions.
The native retry study includes a changed-rule control where preserved scores look
perfect while fresh grading rejects the task; mixed evidence cannot pass acceptance.
Agent judge failures remain missing measurements; opaque grader callback contracts and
observed dependency drift block verified comparisons. These capabilities do not
establish judge accuracy or complete provider accounting.

## Related tools and contributing

Use [pdfhell](https://github.com/multivon-ai/pdfhell) for adversarial document
fixtures, [multivon-mcp](https://github.com/multivon-ai/multivon-mcp) for MCP access,
and [eval-action](https://github.com/multivon-ai/eval-action) for GitHub workflows.
The library can sit alongside your existing tracing system.

Issues and pull requests are welcome. See [CONTRIBUTING.md](CONTRIBUTING.md) for
setup and testing, and [extension contracts](docs/guides/extensions-and-compatibility.mdx)
for custom graders, native environment integrations, report migration and tested
dependency combinations. Apache 2.0 — [Multivon](https://multivon.ai).

Research direction: [regulated enterprise workflows and moat](plans/regulated-enterprise-study.md),
with [public benchmarks for verifier quality](plans/sota-benchmark-program.md).
The [first RAGChecker baseline reproduction](benchmarks/industrial/results/ragchecker-meta-2026-09-17/README.md)
matches published correlations on 280 cases using released predictions and the
unchanged upstream scorer. Our subsequent same-model study above establishes an
initial measurement, not a SoTA claim on those evaluator benchmarks.
