Metadata-Version: 2.5
Name: glovebox-driver
Version: 0.3.1
Summary: Host-side Python driver for a glovebox microVM session
Project-URL: Homepage, https://github.com/AlexanderMattTurner/agent-glovebox
Project-URL: Source, https://github.com/AlexanderMattTurner/agent-glovebox
Author: AlexanderMattTurner
License-Expression: Apache-2.0
Keywords: ai-control,glovebox,microvm,sandbox
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Security
Requires-Python: >=3.11
Requires-Dist: pydantic>=2
Description-Content-Type: text/markdown

# glovebox-driver

The host-side Python that drives one [glovebox](https://github.com/AlexanderMattTurner/agent-glovebox) microVM: boot it, run commands in it, read what it let out, and tear it down. It imports no eval framework, so an integration picks the framework and this package holds the VM.

`inspect-glovebox` and `exploitbench-glovebox` are both built on it.

## What the host needs

The driver runs the `glovebox` command, which drives Docker's `sbx` sandbox runtime. Install glovebox and sign in to `sbx`, then ask whether this host qualifies:

```bash
glovebox sandbox preflight
```

Set `GLOVEBOX_BIN` when the executable is not on `PATH`.

## The modules

| Module       | Owns                                                                       |
| ------------ | -------------------------------------------------------------------------- |
| `run`        | `DisposableRun` — one whole run, from the workspace it mints to its record |
| `cli`        | resolving the `glovebox` executable and running one `sandbox` verb         |
| `config`     | `GloveboxSandboxConfig` — what a VM may be and which hosts it may reach    |
| `session`    | `GloveboxSession` — one live microVM, from boot to teardown                |
| `wedge`      | booting again when the HOST, not the work, wedged the first attempt        |
| `guest_exec` | the argv for one command inside the guest, as the de-privileged user       |
| `evidence`   | reading the policy decision log the sandbox wrote                          |
| `leak`       | reaping a microVM whose owner died before teardown                         |

## Run one

```python
from glovebox_driver import cli
from glovebox_driver.config import GloveboxSandboxConfig
from glovebox_driver.run import EvidenceRequest, disposable_run

cli.preflight()
config = GloveboxSandboxConfig(workspace="/path/to/workspace", boot_timeout=300)
records = EvidenceRequest(traffic="/path/to/traffic.json", refusals="/path/to/refused.json")

with disposable_run(config, evidence=records) as run:
    print(run.exec(["ls", "-la"]).stdout.decode("utf-8"))
```

`disposable_run` owns the whole lifetime. It reaps whatever an earlier run leaked, stages a private copy of the workspace, boots one microVM onto it, and on the way out writes the records you asked for, tears the VM down, and removes what it staged. A run this raises past has already undone every step it took, so a caller that never receives one owes no cleanup.

Leaving `config` out mints an empty workspace under glovebox's shipped allowlist, which is what a transport probe wants.

`run.workspace` is the host directory the boot read, so stage files into it before the boot. After the boot it is a live view on sbx, which binds it, and a stale copy on Kata, which copies it into a disk image mounted elsewhere. Reach the running workspace with `run.exec(...)` at `run.guest_workspace()`.

### Records and how a lost one is reported

Each field of `EvidenceRequest` names a file, and a field left out promises nothing. A record that could not be written is never silent: the close prints what was lost and raises `EvidenceExportFailed`, which carries the `RunRecord` naming the ones that did land. Teardown has already destroyed the disk the record came from, so a run that reports success without it is a run whose audit reads "no red flags" when the truth is "no evidence".

### Booting from a worker thread

A caller that awaits the boot in a worker thread — an `asyncio.to_thread`, an executor — uses `start_run` and `run.close()` instead, and passes a `Handoff`. A cancelled await returns while the boot thread runs on, and the handoff is what makes the microVM it goes on to reach reclaimable by the next run's sweep. The cancelled caller must call `handoff.disown()` on its way out: the boot thread never sees the cancellation, so without that call the entry stays merely armed and no sweep claims it.

`stage_run` is the two-phase form, for a caller that writes into the workspace before the boot. It hands back a `StagedRun` carrying `workspace`. `staged.boot()` consumes that directory, and `staged.discard()` removes it when the caller gives the boot up.

`GloveboxSession.boot` remains the layer under all of this, for a caller that owns its own workspace and teardown. It waits for a ready marker it never creates.

Apache-2.0.
