Metadata-Version: 2.4
Name: nashbench
Version: 0.1.0
Summary: JAX two-player zero-sum imperfect-information games for benchmarking Nash equilibrium solvers.
Keywords: jax,game theory,nash equilibrium,imperfect information,multi-agent reinforcement learning
Author: Eason Yu
Author-email: Eason Yu <breaking0203@gmail.com>
License-Expression: Apache-2.0
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Dist: jax>=0.8
Requires-Python: >=3.12
Project-URL: Documentation, https://easonyu0203.github.io/2p0s-IIG-bench
Project-URL: Source, https://github.com/easonyu0203/2p0s-IIG-bench
Project-URL: Issues, https://github.com/easonyu0203/2p0s-IIG-bench/issues
Description-Content-Type: text/markdown

# nashbench

nashbench is a benchmark for algorithms that learn Nash equilibria in
two-player zero-sum imperfect-information games. Its games are chosen to be
as large as possible while their **exact exploitability** stays feasible to
compute.

-   **Games in JAX**, which run under `jax.jit` and `jax.vmap`. Their rules,
    actions, and observations match
    [OpenSpiel](https://github.com/google-deepmind/open_spiel).
-   **One policy interface**: a function from an observation and its legal
    actions to action probabilities, which plays both seats.
-   **Exact exploitability** of that policy, from milliseconds for poker to
    about 30 seconds on a GPU for games with 10^10 histories.

| Game | Information sets | Terminal histories | Exploitability (GPU) |
| --- | --- | --- | --- |
| Kuhn poker | 12 | 30 | 4 ms |
| Leduc poker | 936 | 5,520 | 4 ms |
| Goofspiel, 6 cards | 23.1 M | 3.7 × 10^8 | 0.6 s |
| Phantom Tic-Tac-Toe | 6.0 M | 9.8 × 10^9 | 18 s |
| Phantom Tic-Tac-Toe, abrupt | 23.3 M | 1.4 × 10^10 | 24 s |
| Dark Hex 3 | 6.1 M | 9.5 × 10^9 | 24 s |
| Dark Hex 3, abrupt | 27.3 M | 1.5 × 10^10 | 32 s |

Times are per call on an NVIDIA RTX 3080 Ti Laptop GPU, after the first call
for a game.

## Install

```sh
uv add nashbench  # Or: pip install nashbench
```

## Quickstart

```python
import jax
import jax.numpy as jnp
import nashbench

game = nashbench.make("leduc_poker")

# A policy maps one player's observation to action probabilities.
params = jax.random.normal(
    jax.random.key(0), (*game.observation_shape, game.num_actions)
)


def policy(observation, legal_action_mask):
    logits = observation @ params  # Your network goes here.
    return jax.nn.softmax(jnp.where(legal_action_mask, logits, -jnp.inf))


# Play one game. The same policy acts for both players.
key = jax.random.key(1)
state, timestep = game.reset(key)
while not timestep.done:
    key, subkey = jax.random.split(key)
    probs = jax.vmap(policy)(timestep.observation, timestep.legal_action_mask)
    action = jax.random.categorical(subkey, jnp.log(probs))  # One per player.
    state, timestep = game.step(state, action)
print(timestep.reward)  # [2]: sums to zero.

# Compute the policy's exact exploitability: 0 at a Nash equilibrium.
print(nashbench.exploitability(game, policy))
```

Before you train, read
[Conventions](https://easonyu0203.github.io/2p0s-IIG-bench/conventions.html).
It explains how players, seats, observations, and rewards work, and how to
play many games in parallel.

## Documentation

See the [documentation](https://easonyu0203.github.io/2p0s-IIG-bench) for
the games, the policy interface, and the API reference. To contribute, see
[CONTRIBUTING.md](https://github.com/easonyu0203/2p0s-IIG-bench/blob/main/CONTRIBUTING.md).

## License

Apache 2.0. See [LICENSE](https://github.com/easonyu0203/2p0s-IIG-bench/blob/main/LICENSE).
