Metadata-Version: 2.5
Name: giul
Version: 0.2.0
Summary: One ruler for what an answer cost in joules, across a GPU fleet.
Project-URL: Homepage, https://github.com/todd427/giul
Author: Todd McCaffrey
License-Expression: Apache-2.0
Keywords: energy,gpu,joules,nvml,telemetry
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3.11
Classifier: Topic :: System :: Monitoring
Requires-Python: >=3.11
Provides-Extra: dev
Requires-Dist: pytest; extra == 'dev'
Requires-Dist: pytest-asyncio; extra == 'dev'
Provides-Extra: nvml
Requires-Dist: nvidia-ml-py>=12.535; extra == 'nvml'
Provides-Extra: remote
Requires-Dist: httpx>=0.24; extra == 'remote'
Description-Content-Type: text/markdown

# giul

*giúl* — Irish for joule. One ruler for "what did this answer cost in energy"
across the FoxxeLabs fleet: Aigne, Tuiscint, Gléas.

Two rulers make a comparison an argument instead of evidence. This is one,
and all three pin it.

```bash
pip install giul              # no dependencies; imports fine with no GPU
pip install giul[nvml]        # + nvidia-ml-py, for the energy counter
pip install giul[remote]      # + httpx, for metering a remote node
```

## What a joule means here

The number a request is charged is **energy above idle**:

```
joules = measured − idle_w × seconds
```

A request pays for the power it *caused*, not for the power the box burns
existing. `idle_w` is supplied by the caller, per node — giul never guesses
it, because a wrong idle floor silently rewrites every figure that node
reports.

## Three rules

1. **Charged above idle.** As above.
2. **No fabricated zeros.** A window that measured ≤ 0 J above idle did not
   measure a free request, it failed to measure one. It degrades to
   `estimated` (with a hint) or `unknown` — never a `sampled` zero.
3. **Estimates are labelled as estimates.** An estimate that reads as a
   measurement is exactly how an efficiency claim stops being evidence. A
   sampler that cannot be reached costs a measurement, never the request.
4. **A shared card is witnessed, not measured.** Board power is per-card and
   nvidia-smi attributes no watts to a PID, so a foreign process's draw lands
   on whatever request is in flight. Such a window records `contended: true`
   and `contended_by`, and is demoted out of `sampled` — the number is real,
   it is simply not yours. *(0.2)*

## Use

```python
from giul import Meter, Node

node = Node(name="iris", gpu_index=0, idle_w=38.0,
            is_local=True, power_endpoint=None)

# async
async with Meter.for_node(node, joules_per_1k_hint=4100.0) as m:
    await retrieve(); m.mark("retrieve")
    await generate(); m.mark("generate")
e = m.result(tokens=412)
# e.joules, e.method, e.backend, e.stages == {"retrieve": Energy, "generate": Energy}

# sync
with Meter.for_node(node, sync=torch.cuda.synchronize).sync() as m:
    ...; m.mark("verify")
e = m.result(tokens=n)
```

`mark(name)` closes the current stage and opens the next. The total is always
computed over the whole window, never by summing stages, so a stage the
sampler was too slow to see cannot corrupt it.

### Contention *(0.2)*

Tell the node what your own serving process looks like, and anything else
resident is a tenant:

```python
node = Node(name="iris", gpu_index=0, idle_w=16.0, is_local=True,
            serving_match="vllm",                    # ours
            contention_ignore=("chrome", "Xorg"),    # a display, not a competitor
            contention_floor_mib=512)                # too small to matter
```

A node that sets no `serving_match` is assumed dedicated and looks for nothing
— 0.1 behaviour, unchanged. Two things learned the hard way and built in:

- **A display is excluded by name, not by size.** The same idle compositor was
  measured at 300 MiB and 594 MiB an hour apart. A floor tuned to exclude it
  today admits it tomorrow, and then every window on that node reads
  `contended` and the measured series is silently empty.
- **Two processes matching `serving_match` is contention.** Another project
  ran vLLM for offline eval and nvidia-smi reported it as `VLLM::EngineCore`,
  character for character what the serving tier reports. When more than one
  is resident, none can be assumed yours.

Checked at the window's two edges, not in the poll loop. A tenant that
arrives and leaves strictly inside one short window is missed — a known gap,
and the honest direction to miss in.

### Probing a card *(0.2)*

```python
from giul import probe_node

state = await probe_node(node)   # None if unreadable — unknown, not fine
# state["memory_free_mib"], state["utilization_pct"], state["compute_apps"]
```

For callers deciding whether to *spend* something on a card. vLLM sizes its
pool as a fraction of total VRAM but refuses to start if that exceeds what is
free, so a wake into a card with a tenant on it does not fail — it
crash-loops until a timeout. `probe_node` lets the caller refuse in
milliseconds with the shortfall named.

**The caller synchronises the GPU before the closing read.** On the counter
backend the register only counts work the card has *finished*; pass
`sync=torch.cuda.synchronize` when the work is local. An HTTP upstream needs
nothing — the completion returning is the sync point.

## Backends

Chosen at runtime in one place (`Meter.for_node`), needing no per-node
configuration. Support is probed once per card and cached.

| `backend` | when | how |
|---|---|---|
| `nvml_counter` | local card, `pynvml` present, counter answers | `nvmlDeviceGetTotalEnergyConsumption` delta — a true measurement of a sub-second window |
| `smi_sampler` | local card, `nvidia-smi` on PATH | integral of `power.draw` samples |
| `remote_agent` | `node.power_endpoint` set | the same integral, sampled by `giul-agent` on that node |
| `none` | otherwise | `estimated` with a hint, else `unknown` |

`method` stays `sampled` / `estimated` / `unknown`. The counter *is* a
measurement, so it reports `method="sampled"`; the distinction lives in
`backend`.

## Tools

```bash
giul-probe                    # what this host's cards support; counter vs sampler
giul-agent --port 9402        # expose a node's power draw; stdlib only, read-only
```

`/power` returns `watts`, `name`, `memory_used_mib`, `memory_total_mib`,
`memory_free_mib`, `utilization_pct`, and `compute_apps` (`pid`, `used_mib`,
`name` per resident process). Every 0.1 field is unchanged; the rest is 0.2.

Bind `giul-agent` to the mesh address, not `0.0.0.0`, unless the node is
otherwise firewalled.

## Not in 0.1

CPU/RAPL, carbon, € cost, storing series, any UI.
