Metadata-Version: 2.4
Name: qqtools
Version: 1.3.4
Summary: A small tool package for qq
Author: qq
License-Expression: MIT
Project-URL: Repository, https://github.com/kzhoa/qqtools.git
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: torch
Requires-Dist: PyYAML
Requires-Dist: matplotlib
Requires-Dist: pandas
Requires-Dist: scipy
Requires-Dist: rich
Provides-Extra: std
Requires-Dist: lmdb; extra == "std"
Requires-Dist: tqdm; extra == "std"
Requires-Dist: requests; extra == "std"
Provides-Extra: plugins
Requires-Dist: rich; extra == "plugins"
Requires-Dist: scikit-learn; extra == "plugins"
Requires-Dist: prompt_toolkit; extra == "plugins"
Provides-Extra: exp
Requires-Dist: libtmux>=0.62.0; extra == "exp"
Requires-Dist: psutil; extra == "exp"
Requires-Dist: nvidia-ml-py; extra == "exp"
Provides-Extra: dev
Requires-Dist: tox; extra == "dev"
Requires-Dist: pytest; extra == "dev"
Requires-Dist: black; extra == "dev"
Requires-Dist: isort; extra == "dev"
Requires-Dist: flake8; extra == "dev"
Requires-Dist: build; extra == "dev"
Provides-Extra: full
Requires-Dist: qqtools[std]; extra == "full"
Requires-Dist: qqtools[plugins]; extra == "full"
Requires-Dist: qqtools[exp]; extra == "full"
Requires-Dist: qqtools[dev]; extra == "full"
Dynamic: license-file

<div style="
  position: relative;
  width: 100%;
  padding-top: 66.66%; 
  margin-bottom: 20px;
  background: #f0f0f0 url('static/banner_960.jpg') center/contain no-repeat;
  background-size: cover;
">
  <img src="static/banner_960.jpg" 
       alt="" 
       style="
         position: absolute;
         top: 0;
         left: 0;
         width: 100%;
         height: 100%;
         opacity: 0;
       ">
</div>

# ✨qqtools✨
[![PyPI Downloads](https://static.pepy.tech/personalized-badge/qqtools?period=total&units=ABBREVIATION&left_color=GREY&right_color=BRIGHTGREEN&left_text=PyPI+Downloads)](https://pepy.tech/projects/qqtools) ![PyPI - Monthly Downloads](https://img.shields.io/pypi/dm/qqtools?color=3cb371&label=Monthly) ![Python version](https://img.shields.io/badge/python->=3.11-blue)

A lightweight library, crafted and battle-tested daily by *qq*, to make PyTorch life a little easier.

It started from the frustration of PyG’s tightly coupled CUDA ecosystem—carefully matching CUDA versions, installing wheel builds from the official index, and repeatedly reinstalling dependencies like `torch-scatter` whenever anything changed. This project brings back a clean, one-line `pip install ...` experience, with no need to worry about CUDA compatibility.

I’ve gathered the repetitive parts of my day-to-day work and refined them into this slim utility library.
It serves as a unified toolkit for handling data, training, and experiments, designed to keep projects moving fast with cleaner code and smoother workflows.

>Built for me, shared for you.

## What it includes

At its core, `qqtools` is a collection of small utilities I use around PyTorch projects:

- data containers such as `qDict` and `qData`
- dataset and dataloader helpers such as `qDictDataset` and `qDictDataloader`
- small neural network helpers such as `qMLP`
- a lightweight training framework, `qpipeline`
- a command-line experiment queue for Linux, `qexp`
- config and serialization helpers for YAML, JSON, pickle, and LMDB

At the core, it is still a practical toolbox for the repetitive parts around experiments.


## Install



```bash
# Core install
pip install qqtools

# Full install
pip install qqtools[full]

# If you only want the experiment queue extras:
pip install qqtools[exp]
```

> While some parts still work with `torch==1.x`, `torch>=2.4` is recommended


## qDict

`qDict` is mainly there for cleaner attribute access in batch-like code:

```python
# Instead of dirty dict brackets:
# batch["input_ids"], batch["attention_mask"]

# Use clean attribute access:
batch = qt.qDict({"input_ids": input_ids, "attention_mask": attention_mask})
out = model(batch.input_ids)
```

## Context scope and `qt.use_ctx`

`qt.ctx` provides a lightweight scoped context. Values set inside `with qt.ctx(...)` are visible only in that scope and its nested calls, and the outer state is restored automatically when the block exits.

Scope exit restores the previous key bindings. If you intentionally mutate a shared mutable object in place through the live context, that mutation is considered caller-managed behavior and may remain visible outside the block.

```python
import qqtools as qt

with qt.ctx(dim=512):
    print(qt.ctx.dim)  # 512

print(qt.ctx.get("dim"))  # None
```

`@qt.use_ctx` is the simplest way to inject context values into a class constructor:

```python
import qqtools as qt


@qt.use_ctx
class AttentionLayer:
    def __init__(self, dim=64, heads=8):
        self.dim = dim
        self.heads = heads


with qt.ctx(dim=512, heads=16):
    layer = AttentionLayer()
    print(layer.dim, layer.heads)  # 512 16
```

Manual constructor arguments still take precedence over injected context values.

## qexp

`qexp` is a lightweight experiment queue for Linux hosts.
It is built around a shared project root, can work on multi-machines with multi-GPUs.

Quick start:

```bash
qexp init --shared-root /mnt/share/myproject/.qexp --machine gpu-a
qexp submit --name demo1 -- python train.py -c config1.yaml
qexp submit --name demo2 -- python train.py -c config2.yaml
qexp submit --name demo3 -- python train.py -c config3.yaml
# 3 tasks will be queued and run sequentially
```

After `init`, `qexp` saves the current `shared_root` and `machine` as CLI context, so you usually do not need to repeat them on every command.

Agent lifecycle is configured separately from how it is launched. `on_demand` (the default)
automatically starts for local submissions and exits after true idleness. `daemon` stays active
until stopped. `qexp agent start` always launches in the background, while `qexp agent run` is
the foreground debugging command.

```bash
qexp init --shared-root /mnt/share/myproject/.qexp --machine gpu-a --agent-mode daemon
qexp agent start
qexp agent status
qexp agent stop
```

### Schema 6 operation

qexp schema 6 uses the Group, Task, and Attempt runtime. Batch-era roots are not
compatible. A drained schema-5 root can be upgraded only when it has no active claim or
running Attempt:

```bash
qexp migrate --shared-root /path/to/project/.qexp --machine gpu1 --to-schema 6
```

The agent owns lease renewal, Recovery, termination, terminal publication, and GPU
reservation release. The runner only starts the training process and writes local process
registration and exit-observation records. Inspect or change the shared lease policy only
while no active claim exists:

```bash
qexp lease-policy show
qexp lease-policy set --ttl-seconds 180 --renew-interval-seconds 10
qexp doctor verify
```

Schema 6 detects clock capability instead of requiring `chronyc` on every host. A qualified
provider permits full bounded-lease coordination; otherwise eligible work runs in holder-bound
local-safe mode and is never expired, remotely recovered, or automatically replaced. `qexp
doctor verify` and `qexp agent status` expose the provider, authority mode, and blocker.

```bash
qexp task share TASK_ID
qexp task share TASK_ID --after 10m --with gpu-b --with gpu-c
qexp task keep-local TASK_ID
qexp task offer TASK_ID --format json
```

`share` is the user-facing control for letting eligible Group workers help while the home
machine remains eligible. `share --after` records a bounded deadline; `keep-local` clears the
shared policy and returns the Task to the home queue. `task offer` is retained for Tasks that
were already submitted with spillover policy and only moves that existing policy into the shared
queue.

For normal task and cleanup workflows:

```bash
qexp submit --group sweep -- python train.py --config a.yaml
qexp batch-submit --group sweep --file runs.yaml
qexp group pause sweep
qexp task retry TASK_ID
qexp task retry TASK_ID --acknowledge-duplicate-risk
qexp clean --task-id TASK_ID --dry-run
qexp clean --older-than-days 30 --limit 100
```

Terminal notifications are disabled by default. Configure the machine-local Feishu Incoming
Webhook from the agent environment (the default, recommended mode):

```bash
qexp config notifications set --enabled
qexp config notifications provider set feishu --enabled \
  --webhook-env QEXP_FEISHU_WEBHOOK --secret-env QEXP_FEISHU_SECRET
export QEXP_FEISHU_WEBHOOK='https://open.feishu.cn/open-apis/bot/v2/hook/...'
export QEXP_FEISHU_SECRET='...'
qexp config notifications show
```

For installations that deliberately accept the shared-root credential risk, a webhook can instead
be persisted under that machine's `.qexp/machines/<machine>/secrets/` directory. The URL is read
from standard input so it does not enter shell history; the explicit acknowledgement is required:

```bash
printf '%s\n' 'https://open.feishu.cn/open-apis/bot/v2/hook/...' |
  qexp config notifications provider set feishu \
    --enabled --credential-source shared_file --webhook-stdin --acknowledge-shared-secret-risk
```

This file is requested as owner-private (`0600`) but remains on the shared control root. Anyone
with access to that storage or its backups may be able to read it. `qexp config notifications show`
never prints the URL. A signing secret, when configured, remains environment-only.

The webhook and secret are read by the process that commits the terminal transition. Non-sensitive
configuration is read at dispatch time, so changes affect future terminal events. Environment
variable value changes require restarting that agent; restarting the agent does not terminate the
running task process. Delivery is synchronous and no-throw with at-most-one send attempt: crashes
or network ambiguity can permanently lose a notification, and qexp does not retry it.

`batch-submit` manifests may set Group workers and nested placement defaults, with per-Task
overrides:

```yaml
group:
  workers: [g1, g2]
defaults:
  placement:
    home_machine: current
    sharing:
      mode: spillover
      fallback_machines: group
tasks:
  - command: [python, train.py]
  - placement:
      sharing:
        mode: private
    command: [python, control.py]
```

During a shared-filesystem outage, the owning agent retains the training process and GPU
reservation in `suspect` and then `isolated` state; it does not create a replacement Attempt
or impose an automatic kill deadline. When shared authority becomes available again, the agent
renews the same claim, recovers the same orphaned Attempt with a new token, or terminates the
old process through its durable termination-decision path if authority changed.

Cleanup waits for required machines to acknowledge removal of matching local GPU reservations,
process manifests, and logs before deleting shared Task and Attempt records. Required machines
are the Task home machine, historical Attempt machines, and the machine that prepared cleanup.
Pending operations report `waiting_ack` and the remaining machine names. Cleanup blocks retry,
claim, cancel, and offer, and its tombstone permanently reserves the Task ID.

`batch-submit` is only a bulk-input command and does not create a public Batch identity.

Python API:

```python
from qqtools.plugins import qexp

task = qexp.submit(
    qexp.load_root_config("/mnt/share/myproject/.qexp", "gpu-a"),
    command=["python", "train.py", "--epochs", "10"],
    name="demo",
)
print(task.task_id)
```

>Note: Run `pip install qqtools[exp]` before use `qexp` command.

## qpipeline

`qpipeline` is a minimal training loop scaffold. It doesn't try to be a heavy framework. You write the project-specific model and task logic, and qpipeline handles the repetitive boilerplate: config-driven startup, train/val loops, metric aggregation, and checkpointing.

A tight training entry:

```python
import torch
from qqtools.plugins.qpipeline import prepare_cmd_args, qPipeline
from qqtools.nn import qMLP

class MyTask:
    def __init__(self, args):
        # Your custom data logic goes here
        self.train_loader, self.val_loader = build_loaders(args)

    def batch_forward(self, model, batch):
        return {"pred": model(batch.x)}

    def batch_loss(self, out, batch):
        loss = torch.nn.functional.mse_loss(out["pred"], batch.y)
        return {"loss": (loss, len(batch.y))}

    def batch_metric(self, out, batch):
        mae = (out["pred"] - batch.y).abs().mean()
        return {"mae": (mae, len(batch.y))}

    def post_metric_to_err(self, result):
        return result["mae"]

class MyPipeline(qPipeline):
    @staticmethod
    def prepare_model(args):
        return qMLP([16, 8, 1])

    @staticmethod
    def prepare_task(args):
        return MyTask(args)

if __name__ == "__main__":
    args = prepare_cmd_args()
    pipe = MyPipeline(args, train=True)
    pipe.fit()
```

Because qpipeline enforces a stable entry contract, it pairs perfectly with qexp for queued execution:

```bash
qexp submit -- python entry.py --config configs/train.yaml
```

For one-off runtime config edits, `qpipeline` also supports dotted CLI overrides after normal
parser handling:

```bash
python entry.py \
  --config configs/train.yaml \
  --task.dataloader.eval_batch_size 32 \
  --task.val_split val_ood \
  --runner.fast_dev_run
```

Configuration follows a standard YAML structure. See [qConfig.md](docs/qConfig_en.md) for details.

## Plugin modules

Under `src/qqtools/plugins/`, there are also:

- `qchem` - tools for reading and processing quantum chemistry outputs
- `qpipeline` - a training pipeline framework built on top of the core torch utilities
- `qhyperconnect` - an implementation of Hyper-Connection for PyTorch

## Test

```bash
tox
```
