Metadata-Version: 2.4
Name: haps-plugin-scheduler-runtime
Version: 0.4.1
Summary: Run HAPS jobs on a French HPC center by submitting them to its scheduler over SSH
Author: Logica Team, Universite de Rennes / IRISA
Author-email: Benjamin De Zordo <benjamin.de-zordo@irisa.fr>
License-Expression: MIT
Project-URL: Homepage, https://www.irisa.fr/equipes/logica
Keywords: haps,hpc,plugin,distributed-computing
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: System Administrators
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: System :: Clustering
Classifier: Topic :: System :: Distributed Computing
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: haps-runtime==0.4.*
Requires-Dist: haps-runtime-executor==0.4.*
Requires-Dist: haps-runtime-remotessh==0.4.*
Requires-Dist: pydantic>=2.10
Provides-Extra: dev
Requires-Dist: ruff>=0.6; extra == "dev"
Dynamic: license-file

# haps-plugin-scheduler-runtime

The maintainer's half of the HAPS scheduler plugin. It runs HAPS jobs on an HPC
centre by submitting them to Slurm over SSH, with **your** account on that
centre — the users who submit the jobs need none.

```bash
pip install haps-plugin-scheduler-runtime
```

## What it does for you

For every job, `SchedulerJob` goes through:

```
connect → setup directories → fetch arguments → fetch ebuffer inputs
        → stage script → execute → push results → push ebuffer outputs
```

It creates a directory per job on the centre, writes the batch script, submits
it, follows it with `squeue` / `sacct`, publishes Slurm's state on the HAPS job,
cancels the Slurm job if the HAPS job is cancelled, and saves Slurm's stdout and
stderr in buffers when the job ends.

What it cannot know is **your application**: what to run, and which files go in
and out. That is the part you write.

## 1. Describe the centre and the resources

```toml
[hpc]
host         = "idris-ssh"                        # an entry of your ~/.ssh/config
user         = "abc1234"
passfile     = "/home/me/.vault/idris"            # chmod 600
workdir_root = "/path/on/the/centre/haps/run"     # written out in full, no $VARIABLE

[resources]
account       = "abc@cpu"
target        = "cpu_p1"                          # the partition
time          = "00:30:00"
nodes         = 1
ntasks        = 4
modules       = ["gromacs/2024.2-mpi"]
poll_interval = 15                                # seconds between two squeue
```

`[hpc]` also takes SSH timeouts; `[resources]` also takes `cpus_per_task`, `qos`,
`environment`, and more. Anything misspelled is refused at startup.

## 2. Write the job

```python
from pathlib import Path
from tempfile import TemporaryDirectory

from haps_plugin_scheduler_runtime import SchedulerJob


class Gromacs(SchedulerJob):
    def script_core(self) -> str:
        """What the batch script runs, after the #SBATCH header and the modules."""
        return f"{self.profile.launch_cmd} gmx_mpi mdrun -s md.tpr -deffnm md"

    def ebinput(self, index: int, name: str, buffer_uuid: str) -> None:
        """One input buffer -> one file in the job's directory on the centre."""
        with TemporaryDirectory() as folder:
            local = Path(folder) / name
            self.client.buffers.get(buffer_uuid).fetch(local)
            self.remotessh.put(local, self.workdir / name)

    def eboutput(self, index: int, name: str, buffer_uuid: str) -> None:
        """One file produced on the centre -> one output buffer."""
        with TemporaryDirectory() as folder:
            self.remotessh.get(self.workdir / name, Path(folder))
            self.client.buffers.get(buffer_uuid).fill(Path(folder) / name)
```

`self.remotessh` talks to the centre, `self.workdir` is this job's own directory
there. Never put what a user sent into `script_core`: it would run under your
account.

## 3. Start the daemon

```python
from haps_plugin_scheduler_runtime import IdrisProfile, SchedulerConfig
from haps_runtime_executor import RuntimeService

config = SchedulerConfig.from_toml("idris.toml", profile=IdrisProfile)
# runtime: your HAPS runtime, found or created with haps_runtime.HapsClient
RuntimeService(runtime, Gromacs, config=config).start()
```

The profile is the centre: `IdrisProfile`, `CinesProfile` or `TgccProfile`
(which speaks `ccc_msub` rather than `sbatch`). The same job class runs on all
three.

## Good to know

- **Dry run** is supported: the script is checked with `--test-only` and
  nothing runs. Your `eboutput` is called as usual, on placeholder files.
- **At the end of a job**, the job's directory on the centre is removed. Override
  `remove_workdir()` to keep it.

Built on `haps-runtime-executor` and `haps-runtime-remotessh`.
