Metadata-Version: 2.4
Name: tasc-hpc-daemon
Version: 0.1.3
Summary: HPC cluster daemon for bridging AI agents to remote compute resources
Author: Tiptree Advanced Systems Corporation, Miles Qi Li
License-Expression: MIT
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: websockets>=14.0
Dynamic: license-file

<p align="center">
  <a href="https://pypi.org/project/tasc-hpc-daemon/"><img src="https://img.shields.io/pypi/v/tasc-hpc-daemon" alt="PyPI version" /></a>
  <a href="https://pypi.org/project/tasc-hpc-daemon/"><img src="https://img.shields.io/badge/Python-3.10%2B-3776AB?logo=python&amp;logoColor=white" alt="Python 3.10+" /></a>
</p>

# HPC Daemon

Connect your HPC cluster or remote server to Althea’s Code Assistant to run commands and submit SLURM jobs.

- **Run commands** on your connected server through Althea’s Code Assistant.
- **Submit SLURM jobs** or run background jobs on servers without SLURM.
- **Choose writable directories** during setup, with operating-system sandbox enforcement where available. See [Safety](#safety) for requirements and limitations.

[Requirements](#requirements) · [Quick start](#quick-start) · [First use](#first-use) · [Commands](#running) · [Safety](#safety)

## Requirements

- Python 3.10+
- A [Tiptree](https://tiptreesystems.com) account
- Linux: Bubblewrap (`bwrap`) with unprivileged user namespaces enabled, to enforce directory restrictions. Setup asks what to do on a host without it.
- macOS: `/usr/bin/sandbox-exec` (Apple Seatbelt)

Supported platforms: Linux and macOS. Native Windows is not supported.

## Quick start

Run these commands on the cluster login node or remote server you want Althea to use:

```bash
pip install tasc-hpc-daemon
hpc-daemon setup --email you@example.com
hpc-daemon start --detach
hpc-daemon status
```

Use the email address of your Althea account. Complete the email verification and choose writable directories when prompted. `--detach` keeps the daemon running after you disconnect; `status` checks the local daemon process.

## First use

Open [Althea](https://althea.tiptreesystems.com) and ask its Code Assistant:

> On my connected server `<hostname>`, show the current directory and whether SLURM is available. Don’t submit a job yet.

Replace `<hostname>` with the daemon ID shown by `hpc-daemon list`. The assistant should run the checks on that server and report the results. If it cannot connect, check `hpc-daemon status` and the [daemon logs](#running).

## Setup

Setup connects to `https://althea.tiptreesystems.com` by default.

The setup wizard:

1. Authenticates via a one-time code sent to your email
2. Presents a **disclaimer** about remote code execution risks
3. Creates an API key for daemon authentication
4. Prompts for an optional **skill** (built-in cluster-specific guidance, or a custom server description)
5. Prompts for **directory restrictions** (where the agent is allowed to write; defaults to `$SCRATCH/tiptree-workspace` or `~/tiptree-workspace`)
6. Prompts for a **job working directory** and optional **server instructions**. The
   default job directory is `hpc_jobs` inside the first approved writable directory.
   Choosing a path outside the approved directories requires explicit confirmation
   because it adds another writable directory.
7. Registers the daemon with the platform

The daemon ID defaults to the machine's hostname. Override with `--daemon-id`.

Re-running setup for the same daemon ID updates the existing profile (no duplicates).

### Non-Interactive Setup

For automated deployments:

```bash
hpc-daemon setup \
    --email you@example.com \
    --no-interview \
    --allowed-dirs ~/workspace ~/scratch \
    --skill mila-hpc \
    --cluster-name my-cluster
```

Use `--cluster-name` to set a human-readable name for the cluster (defaults to the machine's hostname).

> **Warning:** `--no-interview` skips **all** interactive prompts (disclaimer, skill selection, directory restrictions, working directory, server instructions). OTP is still required. With `--allowed-dirs`, job files default to `hpc_jobs` inside the first approved directory. Without `--allowed-dirs`, the agent gets unrestricted filesystem write access.

## Running

```bash
# Start in foreground
hpc-daemon start

# Start in background (detached from the terminal, survives logout)
hpc-daemon start --detach

# Check status
hpc-daemon status

# View logs
tail -f ~/.hpc_daemon/<daemon_id>.log

# Stop
hpc-daemon stop

# List registered daemons
hpc-daemon list
```

If only one daemon profile exists, `--daemon-id` is auto-detected. With multiple profiles, specify it explicitly (e.g., `hpc-daemon start --daemon-id mila-login-1`).

`--detach` forks the daemon into its own session with stdin, stdout and stderr moved off the launching terminal (everything goes to the log file), so it keeps running after the shell or SSH session that started it goes away. Do not background a foreground `hpc-daemon start` from a non-interactive SSH command instead: once that SSH session drops, the daemon's stdout is a pipe nobody reads, and the first log line that fills it blocks the daemon until the server reports it offline.

To force local mode on a SLURM cluster (jobs run as background processes instead of `sbatch`):

```bash
LOCAL_MODE=1 hpc-daemon start
```

## How It Works

Althea’s Code Assistant runs in isolated cloud sandboxes and cannot directly reach machines behind firewalls. The daemon solves this with a reverse-proxy pattern:

1. The daemon runs on your server and opens an **outbound** WebSocket to the Tiptree platform
2. Althea’s Code Assistant also connects **outbound** to the same platform
3. The platform routes messages between them, keyed on your user identity

Because the daemon initiates the connection, no inbound firewall rules are needed.

### Execution Modes

- **PTY mode** (synchronous) — interactive shell sessions for quick commands
- **Job mode** (asynchronous) — batch job submission with automatic wake-on-complete callbacks
  that are persisted locally and retried until delivery succeeds

The daemon auto-detects SLURM. If `sbatch` is available, jobs go through SLURM; otherwise they run as local background processes.

## State

All configuration, job records, and callback delivery state are stored in
`~/.hpc_daemon/state.db` (SQLite). PID files and logs live in the same directory.

## Safety

### Writable directories

When directory restrictions are configured, the daemon launches PTY shells and local jobs inside an operating-system filesystem sandbox. Linux uses Bubblewrap with a read-only view of the host filesystem and writable bind mounts for the configured directories. Host accelerator devices remain available, while `/tmp` and `/dev/shm` are replaced with private temporary storage. macOS uses Apple Seatbelt with write access granted only beneath the configured directories; programs must use `$TMPDIR` rather than a hard-coded `/tmp` path unless `/tmp` was explicitly approved. The daemon state directory remains read-only even when a broader allowed directory contains it.

Advisory Bash function wrappers and redirection checks provide early, readable errors for common file-modifying commands (`rm`, `rmdir`, `mv`, `cp`, `mkdir`, `touch`, `tee`); Bubblewrap or Seatbelt remains the enforcement boundary. Directory navigation is unrestricted, so commands may read outside the writable roots while outside writes remain blocked.

The shell-level checks are advisory and bypassable; the operating-system sandbox is the hard filesystem-write boundary.

### Sandbox requirements

The daemon probes the sandbox before connecting. It fails closed when restrictions cannot be enforced, unless the operator chose during setup to allow running where the sandbox is unavailable, in which case it starts with the advisory checks alone and says so at startup and in its log.

When running without sandbox enforcement, these protections do not hold: the state directory is writable, a session shares the host's process namespace and can signal running jobs, `/tmp` and `/dev/shm` are the host's own, and the recorded process ID of a local job is its launcher rather than a sandbox wrapper.

The daemon removes shell and dynamic-loader startup hooks before invoking trusted launchers, then restores inherited `LD_*` and `DYLD_*` settings only inside the sandbox so cluster CUDA and MPI library paths remain available.

### SLURM jobs

SLURM scripts enter the same sandbox on the compute node, so where enforcement is in effect Bubblewrap must be installed and usable there as well as on the login node.

Node-local scratch is bound writable when `$SLURM_TMPDIR` resolves to an existing directory, so jobs can stage datasets there as usual. When `$SLURM_TMPDIR` resolves to `/tmp`, as it does when SLURM mounts job-private storage there, that disk-backed mount replaces the sandbox's private RAM-backed `/tmp` for the job.

Scratch is skipped when its value would shadow another mount the sandbox relies on: `/`, `/dev/shm`, an approved directory, or a parent of the job working directory. The advisory guardrails accept that directory at run time, but an output redirection to a literal scratch path (rather than one under `$SLURM_TMPDIR`) is still rejected before submission, because the daemon cannot know the path until the job runs.

For restricted SLURM jobs, the daemon streams a daemon-generated wrapper directly to `sbatch`, permits only known single-job scheduling/resource directives, records scheduler and pre-sandbox errors in a protected `.hpc-daemon-control` directory, and creates the payload's stdout/stderr files from inside the sandbox. Slurm arrays are rejected until the daemon can track each task's state and output.

### Local jobs and cancellation

In local mode the sandbox also gives each job its own process namespace, so a PTY session cannot signal a running local job, and the daemon exposes no cancellation for them. The recorded process ID is the outer sandbox process; terminating it leaves the job itself running, so a local job has to be stopped by finding its process on the daemon host.

Job ownership is also enforced: code assistants can only cancel jobs they submitted.

### Shared accounts

> **Shared machines:** Your API key is stored in `~/.hpc_daemon/state.db`. The file and directory are restricted to your user account (`0600`/`0700`), so other users on the same login node cannot read it. However, if multiple people share the same Unix account, they all have access. Do not use the daemon on a shared account.

## License

Licensed under the MIT License. See the `LICENSE` file included in the package.
