Metadata-Version: 2.4
Name: vqflow
Version: 0.2.0
Summary: Versioned VQFlow datasets, model artifacts and managed training from Python.
Author: VQFlow Labs
License-Expression: Apache-2.0
Project-URL: Repository, https://github.com/vqflowlabs/vqflow-python-packages
Classifier: Development Status :: 3 - Alpha
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Operating System :: OS Independent
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# vqflow

Download versioned VQFlow datasets, publish complete model versions and manage
training from Python. When an export version
does not exist, the SDK requests it, waits for generation, downloads its ZIP and
extracts the dataset before returning its directory path.
Both the resource interface and fluent builder use the same export workflow.

This guide describes version `0.2.0`, which requires Python 3.10 or newer.
Install the latest published version from PyPI with:

```sh
python -m pip install --upgrade vqflow
```

For development, install from an authorized repository checkout:

```sh
python -m pip install ./packages/vqflow
```

## Connect

Provide your API key and HTTPS API base URL at runtime. The base URL must include
the API prefix, such as `/api/v1`. The SDK reads only `VQFLOW_API_KEY` and
`VQFLOW_API_URL` when the corresponding constructor arguments are omitted; it
does not load environment files or machine credential stores.

```python
import os
from vqflow import VQFlow

vqf = VQFlow(
    api_key=os.environ["VQFLOW_API_KEY"],
    base_url=os.environ["VQFLOW_API_URL"],
)
```

Use an API key with access to the selected dataset. Downloading an existing
export can use a read-only key. Generating a missing version requires a key
permitted to create exports. The SDK respects the account's existing permissions.

## Resource interface

```python
dataset = (
    vqf.workspace("sample-workspace")
    .project("sample-project")
    .dataset("sample-dataset")
)

dataset_export_path = dataset.export("coco", "training-v1", "./datasets")
```

Use exact workspace, project and dataset names rather than display titles.
Selection is lazy: navigation methods do not issue requests until an export runs.
`dataset_export_path` is a `pathlib.Path` to the extracted dataset directory,
named `vqflow-export-<export-id>` inside the destination directory. For example,
an export with ID 44 saved to `./datasets` returns the absolute path to
`./datasets/vqflow-export-44/`. Its folders and files keep the archive's original
layout. Temporary downloads are removed after extraction. When the server
provides integrity metadata, a verified ZIP is retained in the local cache.

An existing export directory is reused when its full contents verify against
the requested export. Different or incomplete contents require `overwrite=True`.
With that
option, a successful export replaces the entire directory, including any files
you added there; it does not merge old and new dataset contents. Download or
extraction failure preserves the existing dataset and does not return a partial
result.

Version `0.2.0` introduces this directory return value, verified cache and
progress displays. Version `0.1.0` returned the ZIP path. When upgrading to
`0.2.0`, remove any manual unzip step from code that consumes the returned path.

The second `export()` argument is always the **dataset export version**. If
multiple datasets share a name, select the source dataset explicitly:

```python
dataset = (
    vqf.workspace("sample-workspace")
    .project("sample-project")
    .dataset("sample-dataset", dataset_version="source-v2")
)
```

Omitting `dataset_version` requires an unambiguous dataset. Explicit `None`
selects the unversioned dataset. The SDK never silently selects the latest match.

Supported formats are `coco`, `category_folders`, and `vqflow_native`. You can
also use the `ExportFormat` enum. The server supplies a ZIP for every format;
the SDK unpacks it locally. ZIP files within the dataset remain regular files.

## Export builder

```python
from vqflow import DatasetExportBuilder

builder = (
    DatasetExportBuilder(vqf)
    .withWorkspace("sample-workspace")
    .withProject("sample-project")
    .withDataset("sample-dataset")
    .withFormat("coco")
    .withVersion("training-v1")
    .withDestination("./datasets")
)

dataset_export_path = builder.export()
```

The `with...` methods configure the builder. Only `export()` and
`exportCurated()` perform requests. Use `withDatasetVersion(...)` when the source
dataset needs disambiguation; `withVersion(...)` always names the export version.

## Media selection

Exports include all media in the selected dataset unless status filters are
specified. Review and QA filters combine with **AND**:

```python
from vqflow import ReviewStatus, QAStatus

dataset_export_path = (
    builder
    .withVersion("approved-v1")
    .withReview(ReviewStatus.APPROVED)
    .withQA(QAStatus.APPROVED)
    .export()
)
```

The curated shortcut requires both review and QA to be approved:

```python
dataset_export_path = builder.withVersion("curated-v1").exportCurated()
```

`exportCurated()` overrides both filters for that call and leaves the builder's
saved filters unchanged. Each filter supports `APPROVED`, `REJECTED`, `PENDING`,
and `NONE`, as an enum or case-insensitive string. Python `None` removes the
filter; the string `"NONE"` selects media whose actual status is `NONE`.

The resource interface accepts the same filters:

```python
dataset_export_path = dataset.export(
    "coco", "approved-v1", "./datasets", review="APPROVED", qa="APPROVED"
)
```

## Code supplied by VQFlow

The portal's generated code supplies the API address and the exact dataset and
export identities. An exact export ID preserves the stored export's original
media selection, including selections beyond the SDK's review/QA filters:

```python
# Keep your API key private. Do not share it or commit it to source control.
dataset_export_path = dataset.export(
    "coco", "training-v1", "./datasets", export_id=44
)
```

The builder equivalent is `.withExportId(44).export()`. An explicit ID never
creates a replacement if that export is missing, and it cannot be combined with
review/QA overrides or `exportCurated()`. Use a new export version without an
explicit ID when requesting a different selection.

With an explicit export ID, the workspace, project, dataset name and optional
dataset version are checked against the export's saved snapshot. Current
workspace and dataset lists are not consulted, so the code remains usable after
the source dataset moves or its version changes, provided the account still has
access to the stored export. Without an export ID, navigation selects the current
dataset as usual.

Generated code containing an API key is private to you. Do not share the snippet,
include it in a public notebook, or commit it to source control. Revoke a key in
VQFlow if it has been exposed.

## Existing versions and generation

The SDK searches all result pages for the exact dataset and export version. It
selects the requested format and ZIP result, then verifies the saved task's
selection. Other formats may coexist under the same export-version label.

- A ready matching export is downloaded using a fresh access URL.
- A matching queued or running export is awaited and reused.
- An absent export version starts one normal background export.
- A version with incompatible format or filters raises an export conflict.
- Duplicate matching exports and same-name datasets raise an ambiguity error.
- Legacy exports without verifiable selection metadata require a new export
  version. They are not assumed to contain all media.
- A failed or cancelled export raises an error. It is not treated as missing.

Choose a new export version when changing the selection. Existing versions are
reused without rebuilding them from the dataset's current contents. An empty
selection is reported through the normal server export failure.

## Repeated notebook runs and cache controls

For exports with a server-recorded SHA-256 and byte count, repeated calls use
the following checks before returning:

1. Resolve the exact export and request fresh authorized access from VQFlow.
   An expired or revoked API key cannot bypass this check by using local cache.
2. Verify the cached ZIP's complete SHA-256 and length against the server's
   metadata. The cache is scoped to the API address, export ID and ZIP digest;
   rotating download URLs are not used as identities or saved to disk.
3. Read the verified ZIP to reconstruct its file manifest, including each file's
   SHA-256 and size and every directory. Compare it with the stored manifest.
4. Check the entire destination tree: every file's bytes must match, all files
   and directories must exist, and there must be no extra entries or links.

| Local state | Behavior |
| --- | --- |
| Verified ZIP and matching extracted directory | Return the existing directory; no ZIP download or file rewrite. |
| Verified ZIP, missing extracted directory | Extract from cache; no ZIP download. |
| Changed, missing or extra extracted files | Preserve the directory and raise an error; `overwrite=True` permits full replacement from cache. |
| Missing or invalid cache | Download and verify a fresh ZIP; reuse an existing directory if it then verifies. |
| Server reports a different ZIP digest | Use the new artifact identity; never treat an older cached ZIP as a match. |

Verification reads the archive and extracted files on every call. It saves
network transfers and unnecessary writes, but still uses local disk and CPU;
it does not trust timestamps, file sizes alone or an unverified marker file.
These checks establish byte-for-byte integrity at verification time. They do
not assess annotation quality or prevent another process from changing files
after the call returns. Cache reuse requires access to the VQFlow API; there is
no implicit offline mode.

Older exports without server-recorded checksums remain downloadable, but each
call downloads and validates them again. Generate a new export version on a
server that supports integrity metadata to enable verified caching. The SDK
never invents a server checksum from a filename or version label.

Cache controls are the same for resource and builder calls:

```python
# Replace a changed/incomplete dataset using the verified cache when possible.
dataset_export_path = dataset.export(
    "coco", "training-v1", "./datasets", overwrite=True
)

# Fetch fresh bytes even when a valid ZIP is cached. This does not authorize
# replacing different local files; add overwrite=True when that is intended.
dataset_export_path = builder.export(force_download=True)

# Delete cached ZIPs/manifests for one export, keeping extracted directories.
removed_entries = vqf.clear_cache(export_id=44)

# Or clear every cached export for this client's API address.
removed_entries = vqf.clear_cache()
```

`clear_cache()` is local and returns the number of removed export cache entries.
The next export call downloads again. It never deletes returned dataset
directories. Default cache locations are `~/Library/Caches/vqflow` on macOS,
`$XDG_CACHE_HOME/vqflow` or `~/.cache/vqflow` on Linux, and
`%LOCALAPPDATA%/vqflow/Cache` on Windows. Set `VQFlow(..., cache_dir=...)` to use
another directory, for example persistent storage in a hosted notebook.
Keep runtime caches outside source checkouts and package build inputs. They
contain dataset archives and file manifests, but no API keys or signed URLs.
Cache entries remain until explicitly cleared; there is no automatic eviction.
If the cache cannot be written, a successfully verified dataset is still
returned and progress reports `Dataset ready (cache unavailable)`. That call
may need another download next time. Integrity failures are always errors;
cache storage failures never relax ZIP or dataset verification.

## Progress displays

Exports show their current stage while resolving, waiting, verifying,
downloading and extracting. Downloads and extraction show byte progress when
totals are known. Interactive terminals display a spinner/bar; an existing
IPython notebook displays a compact progress card. Redirected output uses
concise stage messages. No optional notebook dependency is required for the SDK.

Use `quiet=True` for scheduled runs or your own presentation:

```python
dataset_export_path = builder.export(quiet=True)
# Also supported by dataset.export(...), dataset.exportCurated(...) and
# builder.exportCurated(...), together with cache and overwrite controls.
```

Progress includes only fixed stage labels and numeric counts, never API keys,
download URLs, dataset file names or raw server responses.

## Waiting, failures and retries

```python
dataset_export_path = dataset.export(
    "coco", "training-v1", "./datasets",
    wait_timeout=7200,
    poll_interval=5,
)
```

The default wait limit is two hours. A timeout does not cancel server work;
`ExportTimeoutError` retains the task/export identifiers. Calling export again
with the same version reuses that task. Scheduled server retries remain active
work even when a previous attempt has an error.

The SDK never automatically repeats an export-creation POST after an uncertain
network outcome. `ExportCreationUncertainError` asks the caller to check the
export version before trying again. The current API does not provide an atomic
get-or-create operation: concurrent clients can create duplicate version
records. The SDK detects observed duplicates and does not silently choose one.

Typed errors are exported from `vqflow`, including `AuthenticationError`,
`PermissionDeniedError`, `NotFoundError`, `AmbiguousResourceError`,
`ExportConflictError`, `ExportFailedError`, `ExportTimeoutError`,
`ExportCreationUncertainError`, `TransportError`, `ProtocolError`, and
`DownloadError`. Errors omit raw server responses and credential-bearing URLs.

Call `vqf.close()` when finished, or use `with VQFlow(...) as vqf:`.

## Credential and download handling

API keys are supplied only to the configured API. ZIP downloads use a separate
unauthenticated connection. HTTPS is required; redirects and ambient proxy
configuration are disabled. Archive data streams to a temporary file and is
expanded into a temporary directory. The dataset becomes available at its final
path only after successful download, extraction and validation. Archive paths
must stay inside that directory; links and unsafe entries are rejected.
Extraction supports stored/deflated ZIP and ZIP64 files, with at most 250,000
archive entries, a 64 MiB central directory and 128 path components. Expanded
data must fit available destination space with a 16 MiB reserve. Files stream
in 1 MiB blocks; there is no fixed total dataset-size or compression-ratio cap.

Never put real keys, passwords, confidential data or captured API responses in
source, examples, tests or package artifacts. Supply credentials only at runtime.

## Exact export handles and local training

The path-returning export methods remain unchanged. For model lineage, resolve
an existing export handle after downloading or generating it:

```python
project = vqf.workspace("sample-workspace").project("sample-project")
dataset = project.dataset("sample-dataset")
dataset_path = dataset.export("coco", "data-v1", "./datasets")
export = dataset.export_version("data-v1", format="coco")

# Or select an exact stored ID without current name resolution.
export = vqf.dataset_export(44)
dataset_path = export.download("./datasets", quiet=True)
```

`export_version()` resolves an existing export and fails if the selection is
missing or ambiguous; it never generates one. The handle binds an exact export
ID, format, dataset identity and version. Its `download()` never silently switches
artifacts. Version labels for datasets, exports and trained models are separate.

Train locally with your preferred library, then upload the checkpoint and the
exact label order used during that training:

```python
model = project.model("sample-detector")
version = model.upload(
    version="weights-v1",
    dataset_export=export,
    weights="./training/checkpoint.pt",
    labels="./training/labels.json",
    framework="pytorch",
    architecture="my-detector",
    model_usage="OBJECT_DETECTION",
    input_signature={"imageChannels": 3},
    preprocessing={"colorOrder": "RGB"},
    dependencies={},
    metrics={"validationAccuracy": 0.91},
    quiet=False,
)
```

The model name is unique within its VQFlow project. Its architecture is separate
metadata; external storage does not require an architecture to be available as
a managed training recipe. `project.model()` is lazy and creates the model only
when uploading or starting training. `vqf.project(22).model("sample-detector")`
uses an exact project ID; `vqf.model(77)` selects an existing model by ID.

Upload requires one checkpoint and one labels file. Additional files may be
supplied as `artifacts={"CONFIG": "./config.json", "METRICS": "./metrics.json"}`.
Metadata arguments `input_signature`, `preprocessing`, `dependencies` and
`metrics` are JSON objects. Record ordered joints, sampling, normalization,
required artifact references and other compatibility details when your model
needs them. The SDK transfers files without loading or executing model code.

Every file is hashed before registration. The server stores an immutable intent;
repeating the same model/version, metadata and files resumes verified 4 MiB
parts. A different intent conflicts instead of replacing a published version.
Only after the server verifies every required artifact does upload return a
`READY` version. An explicit repeated upload may retry its failed or cancelled
receipt. `wait_timeout` and `poll_interval` bound verification waiting; a timeout
keeps the version ID available for reattachment and does not cancel server work.

External origin and metrics are caller-declared provenance. Verifying export
identity and uploaded bytes does not prove that external training used those
data or achieved the supplied metrics. Managed training assigns its own task and
subtask lineage on the server. Registering a model never deploys it to another
application or asks a consumer to report where it is used.

## Retrieve model versions

```python
versions = model.versions()  # Ready versions; include_incomplete=True adds receipts.
version = model.version("weights-v1")
version = vqf.model_version(66)  # Exact ID selection.
files = version.download("./models", quiet=True)
weights_path = files.weights_path
labels_path = files.labels_path
```

`ModelFiles` contains `directory`, `weights_path`, `labels_path`, and optional
`metrics_path` / `config_path`. The version directory is named
`vqflow-model-version-<id>`. Weights are saved as opaque `weights.bin`; this does
not convert their original format. Labels, metrics and config use fixed JSON
filenames. Remote filenames cannot choose local paths. No model deserialization
or framework installation occurs during download.

Each call requests fresh authorized access, verifies the immutable manifest,
then checks the complete local file hashes and directory inventory. Matching
files can be reused without a network artifact transfer. Missing output can be
restored from verified cache. Changed, missing or extra files in an existing
output require `overwrite=True`; replacement publishes the whole verified
version rather than merging directories. `force_download=True` fetches fresh
bytes without implying overwrite. Interrupted or corrupt downloads preserve
prior output and never return a partial version.

```python
files = version.download("./models", overwrite=True, quiet=True)
files = version.download("./models", force_download=True)
vqf.clear_cache(model_version_id=66)
vqf.clear_cache()  # Both dataset and model caches for this client's API.
```

Cache entries contain model bytes, scoped by API, version ID and complete
manifest identity, without API keys or download URLs. Clearing cache preserves
returned model/dataset directories. A destination that independently verifies
against the fresh model manifest remains reusable even after cache clearing;
use `force_download=True` when a new artifact transfer is required.

Use `version.refresh()` or `version.wait(...)` to inspect verification progress.
`version.metadata` exposes compatibility and lineage data without access URLs.
`version.cancel()` and `version.retry()` control external upload receipts;
managed outputs use their training task controls. Write-only upload integrations
can use `vqf.model(model_id).upload(...)`; named navigation needs the corresponding
read access. Referencing an export through the SDK also needs access to that
export. Upload polling uses the write-authorized receipt endpoint.

## Managed training

Read advertised recipes instead of assuming every externally stored model can
be trained by the server:

```python
catalogue = vqf.training_recipes()
recipe = catalogue.recipe("rfdetr-base")
# Inspect recipe["parametersSchema"], recipe["defaults"] and catalogue.compute.
```

Choose compute explicitly and retain a stable run key for each intended training
run. Re-running the same key and settings reattaches to the original task;
changed settings conflict. Compute reservations remain server-enforced across
retries. Starting and retrying training can consume compute resources.

```python
compute = {
    "gpuKey": "your-selected-gpu-key",
    "gpuDisplayName": "Your selected GPU",
    "reservedRuntimeHours": 1,
}
training = model.train(
    recipe=recipe["id"],
    recipe_revision=recipe["revision"],
    version="weights-v2",
    dataset_export=export,
    run_key="sample-detector-experiment-1",
    compute=compute,
    parameters={"epochs": 10},
    include_evaluation=True,
)
training = training.wait(wait_timeout=7200, poll_interval=5, quiet=False)
version = training.model_version(recipe=recipe["id"])
files = version.download("./models")
```

Use actual compute keys and display names advertised for your account. Optional
compute fields follow the server's camelCase contract: `gpuVramGb`, `gpuRamGb`,
`gpuVcpuCount`, `selectedRegion`, `selectedDatacenter`, `selectedCloudType` and
`runpodTemplateExternalId`. No cloud-provider credentials belong in the client.
Parameters are checked against the recipe's advertised flat JSON Schema. Pinning
`recipe_revision` prevents silently using a changed advertised recipe.

```python
training = vqf.training_task(88)
training = training.refresh()
subtasks = training.subtasks()
metrics_by_subtask = training.metrics()
progress_records = training.progress()
training = training.cancel()
training = training.retry()
# Or stop current work and allow rescheduling:
training = training.reschedule()
```

`wait()` distinguishes active scheduled retries from terminal failures. A local
wait timeout leaves server work active; failure/timeout errors retain `task_id`.
Training read permission permits inspection; starting/control needs training
write. Downloading outputs separately requires model artifact read permission.
`training.model_version(version_id=...)` selects an exact output when several
match. The server assigns managed provenance; uploaded external metadata cannot
claim a managed task.

`quiet=True` suppresses model transfer and training progress. Progress shows
fixed stage labels and bounded counts, never local filenames, URLs, keys or raw
server messages. Model errors include `ModelConflictError`, `ModelUploadError`,
`ModelUploadUncertainError`, `ModelVersionFailedError`, and
`ModelVersionTimeoutError`. Training errors include `TrainingFailedError`,
`TrainingTimeoutError`, and `TrainingCreationUncertainError`. An uncertain
creation must be retried with the same version intent or run key.

## License

Licensed under Apache-2.0. The complete license is included in the distribution.
