Metadata-Version: 2.5
Name: fieldtrial
Version: 0.2.0a2
Summary: Find out whether your robot policy actually got better: statistically rigorous real-world evaluation of robot policies.
Project-URL: Homepage, https://github.com/rokbenko/fieldtrial
Project-URL: Repository, https://github.com/rokbenko/fieldtrial
Project-URL: Issues, https://github.com/rokbenko/fieldtrial/issues
Project-URL: Changelog, https://github.com/rokbenko/fieldtrial/blob/main/CHANGELOG.md
Author: Rok Benko
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: confidence intervals,hypothesis testing,lerobot,policy evaluation,robot learning,robotics,statistics,vla
Classifier: Development Status :: 2 - Pre-Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Scientific/Engineering :: Mathematics
Classifier: Typing :: Typed
Requires-Python: >=3.11
Requires-Dist: alembic>=1.13
Requires-Dist: fastapi>=0.115
Requires-Dist: jinja2>=3.1
Requires-Dist: matplotlib>=3.8
Requires-Dist: numpy>=1.26
Requires-Dist: pydantic>=2.7
Requires-Dist: python-multipart>=0.0.9
Requires-Dist: pyyaml>=6.0
Requires-Dist: qrcode>=7.4
Requires-Dist: rich>=13.8
Requires-Dist: scipy>=1.13
Requires-Dist: sqlalchemy>=2.0.30
Requires-Dist: sse-starlette>=2.1
Requires-Dist: typer>=0.15
Requires-Dist: uvicorn[standard]>=0.30
Provides-Extra: dev
Requires-Dist: httpx>=0.28; extra == 'dev'
Requires-Dist: hypothesis>=6.130; extra == 'dev'
Requires-Dist: import-linter>=2.1; extra == 'dev'
Requires-Dist: mypy>=2.0; extra == 'dev'
Requires-Dist: pre-commit>=4.0; extra == 'dev'
Requires-Dist: pytest-cov>=6.0; extra == 'dev'
Requires-Dist: pytest>=8.3; extra == 'dev'
Requires-Dist: ruff>=0.16; extra == 'dev'
Requires-Dist: scipy-stubs>=1.15; extra == 'dev'
Requires-Dist: statsmodels>=0.14.4; extra == 'dev'
Requires-Dist: types-pyyaml>=6.0; extra == 'dev'
Provides-Extra: docs
Requires-Dist: mkdocs-material>=9.6; extra == 'docs'
Requires-Dist: mkdocs<2,>=1.6; extra == 'docs'
Requires-Dist: mkdocstrings[python]>=0.29; extra == 'docs'
Provides-Extra: openpi
Requires-Dist: websockets>=13; extra == 'openpi'
Description-Content-Type: text/markdown

# fieldtrial

**Find out whether your robot policy actually got better.**

[![CI](https://github.com/rokbenko/fieldtrial/actions/workflows/ci.yml/badge.svg)](https://github.com/rokbenko/fieldtrial/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/fieldtrial.svg)](https://pypi.org/project/fieldtrial/)
[![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](https://github.com/rokbenko/fieldtrial/blob/main/LICENSE)

fieldtrial is an open-source Python framework for statistically rigorous, real-world
evaluation of robot policies. LeRobot trains the policy; fieldtrial tells you whether it
actually got better.

## Why

Real-world evaluation is slow, noisy and ad hoc. With 40 rollouts, the 95% confidence
interval for a success rate is 12 to 15 percentage points wide on each side, so many apparent
improvements from a fine-tuning run are just noise. fieldtrial helps at every stage:

- **Before:** how many rollouts do I need? What is the smallest difference this
  evaluation can detect?
- **During:** a randomized, blinded schedule, and a phone-friendly operator console that
  records the outcome, the furthest stage reached and the failure mode of each rollout.
- **After:** the right statistics for the design (exact tests, paired analyses,
  multiplicity control), honest wording, and a self-contained report.

## 30-second demo

```console
$ uvx fieldtrial compare 74/80 91/120
Arm 1: 74/80 = 92.5% (95% CI 84.6%–96.5%)
Arm 2: 91/120 = 75.8% (95% CI 67.4%–82.6%)
Difference +16.7 pp, Newcombe 95% CI [+6.2 pp, +26.0 pp]
Boschloo p = 0.0020 (two-sided); rejects H0 at α = 0.05
Fisher p = 0.0022

$ uvx fieldtrial power --p1 0.76 --p2 0.90      # rollouts needed to detect 76% → 90%
112 per arm (pooled-z)

$ uvx fieldtrial demo                           # a simulated study in the console
```

`fieldtrial demo` opens a half-run, blinded study with a simulated robot. Run the rest
from the keyboard (Space, Space, Enter), unblind, and read the report.

<p>
  <img src="https://raw.githubusercontent.com/rokbenko/fieldtrial/main/docs/assets/console-running.png" alt="The operator console on a phone: a trial in progress with its blind code, timer and a large Stop button" width="260">
  <img src="https://raw.githubusercontent.com/rokbenko/fieldtrial/main/docs/assets/console-label.png" alt="Labelling a trial: furthest stage reached, why it ended, failure tags" width="260">
</p>
<p>
  <img src="https://raw.githubusercontent.com/rokbenko/fieldtrial/main/docs/assets/report.png" alt="The HTML report: summary, success rate per arm with confidence intervals, primary analysis and a forest plot" width="560">
</p>

<!-- A short GIF of a trial run in the console goes here. -->

## Quickstart

```console
$ uv tool install fieldtrial              # or: pip install fieldtrial
$ fieldtrial init my-study                # writes my-study/study.yaml
$ fieldtrial plan my-study --baseline 0.75
$ fieldtrial lock my-study                # freezes the design and randomizes the schedule
$ fieldtrial serve my-study --lan         # scan the QR code with a phone at the robot
$ fieldtrial unblind my-study
$ fieldtrial report my-study              # my-study/reports/report.html
```

A study is one folder: `study.yaml` (the design) and a SQLite database. Everything is
local and works offline; fieldtrial sends no telemetry.

- [Documentation](https://rokbenko.github.io/fieldtrial/): quickstart, concepts, guides
  and a statistics reference.
- [Re-analysis of Dream Machines' published pi0.5 results](https://github.com/rokbenko/fieldtrial/tree/main/examples/dream-machines-pi05):
  which of 30 published comparisons the data actually resolve.

## What you get

- **Design:** randomized complete blocks with balanced arm order, blind codes, a design
  hash, locking and logged amendments.
- **Console:** sessions with a rig checklist, a timer, stage and failure-tag labels,
  10-second undo, invalid trials with automatic rescheduling, a live mirror screen,
  keyboard and foot-pedal keys, LAN access with a QR code.
- **Analysis:** the primary test follows from the locked design (exact McNemar with a
  Tango interval, Cochran–Mantel–Haenszel, or Cochran's Q with Holm), plus an
  independent-samples sensitivity analysis, stage funnels, time to success, drift checks
  and a list of every deviation from the plan.
- **Reports:** self-contained HTML with charts, Markdown for pull requests, and a
  versioned JSON results model.
- **Integration:** a REST API with a dependency-free Python client for custom runtimes,
  CSV import and export, and `fieldtrial.stats` as a library.

Every statistical function is tested against an independent reference implementation.
Reports only describe a difference when the pre-registered test rejects; otherwise they
say what the study could have detected.

## How it fits with LeRobot and openpi

fieldtrial complements LeRobot and openpi and never forks them. It is not a training
framework, a simulation benchmark, a robot driver, a labeling platform or a cloud
service. In manual mode, fieldtrial schedules and records the trials and you run the
robot however you like. For real blinding, the `command` runner launches your rollout
command (for example `lerobot-rollout`) for each arm, and the `openpi_router` runner
routes an openpi client's traffic to each trial's policy server; see
[Real blinding with runners](https://rokbenko.github.io/fieldtrial/guides/runners/).

## How to cite

If fieldtrial helps your research, please cite it (see
[CITATION.cff](https://github.com/rokbenko/fieldtrial/blob/main/CITATION.cff)):

```bibtex
@software{benko_fieldtrial,
  author  = {Benko, Rok},
  title   = {fieldtrial: statistically rigorous real-world evaluation for robot policies},
  url     = {https://github.com/rokbenko/fieldtrial},
  license = {Apache-2.0}
}
```

For evaluation practice in general, see Kress-Gazit et al. (2024), *Robot Learning as an
Empirical Science: Best Practices for Policy Evaluation*, arXiv:2409.09491.

## Development

See [CONTRIBUTING.md](https://github.com/rokbenko/fieldtrial/blob/main/CONTRIBUTING.md).
