Metadata-Version: 2.4
Name: magma_bench
Version: 2.0.0
Summary: Evaluate agents on interactive robotic tasks with MAGMA
Project-URL: Documentation, https://magma-rob.github.io/docs/intro
Project-URL: Repository, https://github.com/MAGMA-rob/magma-bench
Project-URL: Issues, https://github.com/MAGMA-rob/magma-bench/issues
Project-URL: Changelog, https://github.com/MAGMA-rob/magma-bench/blob/main/CHANGELOG.md
Classifier: Development Status :: 5 - Production/Stable
Requires-Python: <3.13,>=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: magma_core[simulation]<3.0.0,>=2.0.0
Requires-Dist: magma_scenarios<3.0.0,>=2.1.0
Requires-Dist: packaging<27,>=24
Requires-Dist: pydantic<3,>=2.4
Requires-Dist: PyYAML<7,>=6.0
Requires-Dist: requests<3,>=2.34.2
Requires-Dist: torch<3,>=2.8
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Provides-Extra: video
Requires-Dist: imageio-ffmpeg>=0.5; extra == "video"
Requires-Dist: numpy>=1.23; extra == "video"
Requires-Dist: Pillow>=10; extra == "video"
Dynamic: license-file

# MAGMA-BENCH

Evaluate agents on interactive robotic tasks with MAGMA. MAGMA-BENCH loads
compiled benchmark episodes, runs them in simulation, communicates with an agent
through the MAGMA HTTP protocol, and saves detailed results and metrics.

**Version 2.0.0 is the stable v2 release.** It requires MAGMA Core 2.0.0
and MAGMA Scenarios 2.1.0 or later within their respective v2 release series.

[Official documentation](https://magma-rob.github.io/docs/intro) ·
[Report an issue](https://github.com/MAGMA-rob/magma-bench/issues)

## Installation

Python 3.12 is required. Simulation installation has been checked on Linux
x86_64. GPU simulation and rendering require compatible system drivers.

```bash
python -m pip install "magma_bench==2.0.0"
```

This installs `magma_core[simulation]>=2.0.0,<3.0.0` and
`magma_scenarios>=2.1.0,<3.0.0` with their Python dependencies. The benchmark
files and the agent server are separate inputs and are not bundled with this
package.

To upgrade to the latest stable release, use:

```bash
python -m pip install --upgrade magma_bench
```

## Usage

Start a compatible MAGMA agent server, then run a compiled benchmark directory:

```bash
magma-bench run \
  --benchmark-root /path/to/compiled-benchmark \
  --results-path /path/to/results/experiment-1 \
  --agent-address http://127.0.0.1:8888 \
  --run-name experiment-1
```

The agent must expose `/health`, `/v1/info`, and `/v1/responses` using protocol
version 2.0. Configuration can be supplied with `--config-path`; command-line
options override its benchmark settings.

Useful options include:

```bash
magma-bench run --help
magma-bench run --benchmark-root /path/to/benchmark --results-path ./eval/coffee --scenarios coffee_comp
magma-bench run --benchmark-root /path/to/benchmark --results-path ./eval/no-judge --skip-judge
magma-bench run --benchmark-root /path/to/benchmark --results-path ./eval/logged --model-logs --videos all
magma-bench run --benchmark-root /path/to/benchmark --results-path ./eval/failures --videos planner-failure
```

Video output uses the optional dependencies:

```bash
python -m pip install "magma_bench[video]==2.0.0"
```

Merge all completed parts of one evaluation into a YAML analysis:

```bash
magma-bench stats ./stats.yaml ./eval/model-name /path/to/compiled-benchmark
```

The results directory must contain one subdirectory per part, each with `run.json`
and `result.json`. All compiled episodes must be covered exactly once. Partial
runs and infrastructure failures stop the command without writing the YAML.
The `name` field is the results directory name. The file includes success rates,
conditional robustness, failure statuses, and scenario, skeleton, semantic,
length, and track breakdowns. Stage progression is the mean of each episode's
fraction of completed stages, using stage counts from the compiled benchmark.
A benchmark fingerprint mismatch is reported in the YAML and on stderr because
episode IDs alone cannot certify that compiled contents match exactly.

Version 2 uses result schema `3.0`. Terminal statuses are `success`,
`invalid_format`, `textual_hallucination`, `non_requested_action`,
`non_physical_exceeded`, `judge_failure`, `env_failure`, `budget_exceeded`,
and `infrastructure_failure`. An infrastructure failure returns the episode to
the end of its current simulator group, for at most three attempts per launch.
The progress log records every attempt; episode results, model logs, and videos
retain the final attempt. Resume a compatible run with `--results-path`; a
resumed infrastructure failure receives a fresh retry allowance. Result schema
`2.0` and older agent adapters are unsupported.

See the [official documentation](https://magma-rob.github.io/docs/intro) for
benchmark preparation, configuration, agent setup, metrics, and result formats.

## Install from source

Install the tagged release from GitHub:

```bash
python -m pip install "magma_bench @ git+https://github.com/MAGMA-rob/magma-bench.git@v2.0.0"
```

Dependencies are resolved from PyPI. For local development, clone the repository
and run `python -m pip install -e ".[dev,video]"`.

## License

[BSD 2-Clause](https://github.com/MAGMA-rob/magma-bench/blob/main/LICENSE).
