Metadata-Version: 2.5
Name: fiduciarybench
Version: 0.1.2
Summary: FiduciaryBench, by WealthSchema: an open benchmark of AI behavior in regulated wealth-management work (Reg BI suitability, RMD and wash-sale mechanics), with provenance-gated answer keys derived from primary sources. Runners, deterministic scoring, and report output.
Project-URL: Homepage, https://www.wealthschema.com/benchmark
Project-URL: Repository, https://github.com/CapsteraSupport/wealthschema
Author: WealthSchema
License: MIT
Keywords: ai-agent-evaluation,benchmark,eval,fiduciarybench,finra,llm-evaluation,reg-bi,rmd,sec,suitability,wash-sale,wealth-management,wealthschema
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Software Development :: Testing
Requires-Python: >=3.9
Description-Content-Type: text/markdown

# fiduciarybench

**FiduciaryBench, by [WealthSchema](https://www.wealthschema.com)** — an open
benchmark of AI behavior in regulated wealth-management work: Reg BI
suitability, required-minimum-distribution mechanics, and wash-sale
mechanics.

Every answer key is **computed by code or derived from a quoted
primary-source passage** (SEC, FINRA, IRS/Treasury, U.S. Code) that
exact-matches a locally held, versioned corpus. Open-web knowledge is
inadmissible by construction. Full methodology, item provenance, and the
standing challenge policy: https://www.wealthschema.com/benchmark

## Install

```bash
pip install fiduciarybench
```

No dependencies. Python 3.9+.

## Run a model against the public split

```bash
export GEMINI_API_KEY=…            # or ANTHROPIC_API_KEY / OPENAI_API_KEY / DEEPSEEK_API_KEY
fiduciarybench run \
  --provider gemini --model gemini-2.5-pro \
  --items fiduciarybench-public.json \
  --out report.json
```

Providers: `anthropic`, `openai`, `gemini`, `deepseek`, `mistral`, `together`. Runs request
temperature 0; model families that reject the temperature parameter
outright (the Claude 5 family, GPT-5.x reasoning models) run at provider
default, and the report's `sampling` / `sampling_notes` fields say so —
the sampling regime of every run is on the record. Scoring is
deterministic either way: computed items score within the item's stated
tolerance; derived items require the exact option id. A response the
scorer cannot parse counts as incorrect and is tallied as a parse failure
in the report.

## Score your own transcripts

For a model with no runner here (or an in-house system), produce a JSON
object of `item id → response text` and score it offline:

```bash
fiduciarybench score --items fiduciarybench-public.json \
  --responses responses.json --model-label my-system --out report.json
```

## As a library

```python
from fiduciarybench import load_items, score_item, aggregate

items = load_items("fiduciarybench-public.json")
results = [score_item(item, my_model(item)) for item in items]
print(aggregate(results))
```

## What the public split is (and is not)

The public split is the open, citable subset. A held-out split of the
same families backs private evaluation runs and anti-gaming rotation and
never ships. Training on the public split is not license-restricted, but
a score on a memorized split is a score on memory — the held-out split
exists to catch exactly that.

## Challenge policy

Every item carries its provenance: cited passages, derivation or compute
id, verifier record, version history. If you can demonstrate an answer
key is wrong, the erratum is published conspicuously and credited to
you. benchmark@wealthschema.com
