Metadata-Version: 2.4
Name: dataerai-sdk
Version: 0.2.0b35
Summary: Python SDK for the Dataerai transfer daemon
Author: Dataerai
License: DATAERAI PROPRIETARY SOFTWARE LICENSE AGREEMENT
        
        Version 1.2
        
        Copyright (c) 2026 Dataerai, Inc. All Rights Reserved.
        
        ================================================================================
        NOTICE: THE SOFTWARE IS PROPRIETARY AND CONFIDENTIAL. NO RIGHTS ARE GRANTED
        EXCEPT AS EXPRESSLY AND NARROWLY SET FORTH BELOW. ALL RIGHTS NOT EXPRESSLY
        GRANTED ARE RESERVED BY DATAERAI, INC. THIS AGREEMENT APPLIES EQUALLY TO USE
        BY HUMANS AND BY ARTIFICIAL INTELLIGENCE AGENTS (SEE SECTION 4). IF YOU DO NOT
        AGREE TO EVERY TERM OF THIS AGREEMENT, YOU HAVE NO LICENSE AND MUST NOT ACCESS,
        INSTALL, COPY, OR USE THE SOFTWARE IN ANY MANNER.
        ================================================================================
        
        This Dataerai Proprietary Software License Agreement (this "Agreement") is a
        binding legal agreement between Dataerai, Inc., a corporation organized under
        the laws of the Commonwealth of Pennsylvania ("Licensor"), and the single
        identified individual or legal entity that has been expressly authorized in
        writing by Licensor to receive the Software ("Licensee"). By accessing,
        installing, copying, or using the Software, or by clicking to accept, Licensee
        agrees to be bound by this Agreement.
        
        --------------------------------------------------------------------------------
        1. DEFINITIONS
        --------------------------------------------------------------------------------
        
        1.1 "Software" means the Dataerai software package, in any and all forms,
            including without limitation source code, object code, byte code, binaries,
            scripts, schemas, configuration, models, weights, prompts, embeddings,
            graph taxonomies, registries, metadata, libraries, application programming
            interfaces, command-line tooling, and any associated materials, together
            with all updates, upgrades, patches, modifications, enhancements, and
            Documentation, whether delivered now or in the future.
        
        1.2 "Documentation" means any technical or user materials, in any medium,
            provided or made available by Licensor relating to the Software.
        
        1.3 "Authorized Purpose" means solely the internal, non-production evaluation
            of the Software by Licensee, on systems owned and controlled by Licensee,
            strictly within the scope, seat count, term, and field of use expressly
            stated in a written authorization signed by an officer of Licensor. Absent
            such a written authorization, the Authorized Purpose is null and no use is
            permitted.
        
        1.4 "Confidential Information" means the Software and any non-public information
            disclosed by Licensor, whether or not marked confidential. The Software is
            deemed Confidential Information in its entirety.
        
        --------------------------------------------------------------------------------
        2. GRANT OF LICENSE
        --------------------------------------------------------------------------------
        
        2.1 Subject to Licensee's strict and continuous compliance with every term of
            this Agreement, Licensor grants to Licensee a personal, non-exclusive,
            non-transferable, non-sublicensable, non-assignable, revocable, royalty-
            bearing-or-evaluation-only, and fully revocable license to use the Software
            solely for the Authorized Purpose and solely during the Term.
        
        2.2 The license granted in Section 2.1 is the entirety of the rights granted.
            It conveys no right to copy (except a single archival copy as required for
            backup, retaining all notices), no right to modify, no right to create
            derivative works, no right to distribute, no right to sublicense, no right
            to display or perform publicly, and no right to use in production, in a
            service bureau, or on behalf of any third party.
        
        2.3 No license, immunity, or other right is granted by implication, estoppel,
            exhaustion, or otherwise. No patent, trademark, or other intellectual
            property right is licensed except the limited use right in Section 2.1.
        
        --------------------------------------------------------------------------------
        3. RESTRICTIONS
        --------------------------------------------------------------------------------
        
        Licensee shall not, and shall not permit, enable, or assist any third party or
        AI Agent (as defined in Section 4) to:
        
        3.1 copy, reproduce, republish, upload, post, transmit, or otherwise duplicate
            the Software except for the single archival copy permitted in Section 2.2;
        
        3.2 modify, adapt, translate, port, fork, or create derivative works of the
            Software, in whole or in part;
        
        3.3 sell, resell, rent, lease, lend, distribute, transfer, disclose, host,
            sublicense, time-share, or otherwise make the Software available to any
            third party, whether for value or not;
        
        3.4 reverse engineer, decompile, disassemble, decrypt, extract, or otherwise
            attempt to derive or reconstruct the source code, architecture, algorithms,
            model weights, training data, graph taxonomies, or trade secrets embodied
            in the Software, except, and only to the minimum extent, that this
            prohibition is unenforceable under applicable law and only after written
            notice to Licensor;
        
        3.5 use the Software to develop, train, benchmark, or improve any competing,
            similar, or alternative product, model, or service;
        
        3.6 conduct, publish, or disclose any benchmark, performance, security, or
            comparative test or analysis of the Software without Licensor's prior
            written consent;
        
        3.7 remove, alter, obscure, or fail to reproduce any copyright, proprietary,
            confidentiality, or other notice contained in or on the Software;
        
        3.8 circumvent, disable, or interfere with any license-key, telemetry, usage-
            metering, digital-rights-management, or access-control mechanism;
        
        3.9 use the Software in excess of the authorized seat count, scope, field of
            use, or Term, or for any purpose other than the Authorized Purpose; or
        
        3.10 use the Software in violation of any applicable law, regulation, or third-
             party right, including export control and sanctions laws.
        
        --------------------------------------------------------------------------------
        4. ARTIFICIAL INTELLIGENCE, AGENTS, AND AUTOMATED USE
        --------------------------------------------------------------------------------
        
        4.1 Definition. "AI Agent" means any artificial intelligence or machine-
            learning system, model, or software — including any autonomous or semi-
            autonomous agent, large language model, multi-agent system, robotic process
            automation, bot, crawler, scraper, copilot, or other automated process —
            that accesses, operates, invokes, queries, or acts upon the Software,
            whether or not under direct human supervision, and whether operating on
            Licensee's behalf, on Licensee's systems, or using Licensee's credentials,
            tokens, API keys, or sessions.
        
        4.2 Agents Bound by this Agreement. This Agreement applies in full to any access
            to or use of the Software by, through, or on behalf of an AI Agent. Every
            restriction, obligation, and prohibition that binds Licensee binds equally
            any AI Agent that Licensee deploys, authorizes, instructs, integrates, or
            permits to interact with the Software. An AI Agent is not a separate,
            independent, or exempt user, and acquires no rights of its own under this
            Agreement.
        
        4.3 Full Responsibility for Agent Conduct. Licensee is fully and solely
            responsible for all acts and omissions of any AI Agent acting on Licensee's
            behalf, on Licensee's systems, or via Licensee's credentials, as if those
            acts were Licensee's own. The autonomy of the AI Agent, the absence of human
            supervision, or the absence of specific human intent is not a defense to any
            breach. Licensee shall not use an AI Agent to do, attempt, or enable
            anything that Licensee is itself prohibited from doing under this Agreement.
        
        4.4 No Training, Ingestion, or Derivation. Licensee shall not, and shall not
            permit any AI Agent or third party to, use the Software or any portion of it
            — including its source code, object code, structure, schemas, graph
            taxonomies, registries, metadata, model weights, prompts, embeddings,
            Documentation, or Outputs — as input to, as training, fine-tuning, or
            alignment data for, as a retrieval or context corpus for, or as a basis to
            distill, replicate, or derive, any AI Agent, model, dataset, or competing
            system.
        
        4.5 No Automated Extraction. Licensee shall not use any AI Agent to scrape,
            crawl, harvest, index, reverse engineer, or otherwise extract the Software,
            its structure, its parameters, or its Outputs, including by means of
            automated or adversarial querying intended to reconstruct or approximate the
            Software's behavior, weights, taxonomies, or underlying data.
        
        4.6 Outputs. "Outputs" means any data, content, predictions, classifications,
            embeddings, analyses, graphs, or other results generated by the Software or
            by any AI Agent through use of the Software. Outputs constitute Confidential
            Information, are licensed (not assigned) to Licensee solely for the
            Authorized Purpose, and remain subject to every restriction in this
            Agreement. Licensee obtains no right to use Outputs to train, evaluate, or
            improve any AI Agent, model, or competing system.
        
        4.7 Disclosure of Agentic Use. Upon Licensor's request, Licensee shall disclose
            whether, and the manner in which, any AI Agent has accessed or used the
            Software, and shall cooperate with any audit under Section 8 directed at
            such use.
        
        --------------------------------------------------------------------------------
        5. OWNERSHIP
        --------------------------------------------------------------------------------
        
        5.1 The Software is licensed, not sold. Licensor and its licensors retain all
            right, title, and interest in and to the Software and all intellectual
            property rights therein. Licensee acquires no ownership interest of any
            kind. All feedback, suggestions, and ideas provided by Licensee — whether
            authored by a human or generated by an AI Agent — are hereby assigned to
            Licensor without restriction or compensation.
        
        --------------------------------------------------------------------------------
        6. CONFIDENTIALITY
        --------------------------------------------------------------------------------
        
        6.1 Licensee shall hold the Software and all Confidential Information in strict
            confidence, shall not disclose it to any person or AI Agent other than
            Licensee's employees with a need to know who are bound by written
            obligations at least as protective as this Agreement, and shall use it
            solely for the Authorized Purpose. Licensee shall protect it using no less
            than the degree of care it uses for its own most sensitive information, and
            in no event less than a reasonable degree of care. Licensee shall not input,
            expose, or transmit the Software or Confidential Information to any third-
            party AI Agent or service that would acquire rights in, or use for training,
            the material so disclosed.
        
        --------------------------------------------------------------------------------
        7. TERM AND TERMINATION
        --------------------------------------------------------------------------------
        
        7.1 This Agreement and the license commence on the date Licensee first accesses
            the Software and continue only for the period expressly authorized in
            writing by Licensor (the "Term"). If no period is stated, the Term is
            thirty (30) days.
        
        7.2 Licensor may terminate or suspend this Agreement and the license at any
            time, for any reason or no reason, with or without notice, in its sole
            discretion. This license is revocable at will.
        
        7.3 This Agreement terminates automatically and immediately upon any breach by
            Licensee, without notice and without opportunity to cure.
        
        7.4 Upon any expiration or termination, all rights granted cease immediately,
            and Licensee shall, within five (5) days, cease all use, permanently delete
            or destroy all copies of the Software (including the archival copy, all
            Outputs, and all derivatives and extracts), purge the Software and Outputs
            from any AI Agent context, cache, or store, and certify such destruction in
            writing to Licensor upon request. Sections 1, 3, 4, 5, 6, 7.4, and 8
            through 13 survive termination.
        
        --------------------------------------------------------------------------------
        8. AUDIT
        --------------------------------------------------------------------------------
        
        8.1 Licensor may, upon reasonable notice, audit Licensee's use of the Software,
            including by inspecting records and systems and reviewing AI Agent logs and
            access histories, to verify compliance. Licensee shall cooperate. Any
            unauthorized use revealed by an audit constitutes a material breach.
        
        --------------------------------------------------------------------------------
        9. DISCLAIMER OF WARRANTIES
        --------------------------------------------------------------------------------
        
        9.1 THE SOFTWARE IS PROVIDED "AS IS" AND "AS AVAILABLE," WITH ALL FAULTS AND
            WITHOUT WARRANTY OF ANY KIND. TO THE MAXIMUM EXTENT PERMITTED BY LAW,
            LICENSOR DISCLAIMS ALL WARRANTIES, WHETHER EXPRESS, IMPLIED, STATUTORY, OR
            OTHERWISE, INCLUDING ANY IMPLIED WARRANTIES OF MERCHANTABILITY, FITNESS FOR
            A PARTICULAR PURPOSE, TITLE, ACCURACY, AND NON-INFRINGEMENT, AND ANY
            WARRANTY ARISING FROM COURSE OF DEALING OR USAGE OF TRADE. LICENSOR DOES NOT
            WARRANT THAT THE SOFTWARE WILL BE UNINTERRUPTED, ERROR-FREE, OR SECURE, OR
            THAT ANY OUTPUT GENERATED BY THE SOFTWARE OR BY ANY AI AGENT WILL BE
            ACCURATE, COMPLETE, OR FIT FOR ANY PURPOSE.
        
        --------------------------------------------------------------------------------
        10. LIMITATION OF LIABILITY
        --------------------------------------------------------------------------------
        
        10.1 TO THE MAXIMUM EXTENT PERMITTED BY LAW, LICENSOR SHALL NOT BE LIABLE FOR
             ANY INDIRECT, INCIDENTAL, SPECIAL, CONSEQUENTIAL, EXEMPLARY, OR PUNITIVE
             DAMAGES, OR FOR ANY LOSS OF PROFITS, REVENUE, DATA, OR GOODWILL, ARISING
             OUT OF OR RELATING TO THIS AGREEMENT, THE SOFTWARE, OR ANY ACTION TAKEN BY
             AN AI AGENT, REGARDLESS OF THE THEORY OF LIABILITY AND EVEN IF ADVISED OF
             THE POSSIBILITY OF SUCH DAMAGES.
        
        10.2 LICENSOR'S TOTAL CUMULATIVE LIABILITY ARISING OUT OF OR RELATING TO THIS
             AGREEMENT SHALL NOT EXCEED THE GREATER OF (A) THE FEES ACTUALLY PAID BY
             LICENSEE TO LICENSOR FOR THE SOFTWARE IN THE THREE (3) MONTHS PRECEDING THE
             CLAIM, OR (B) ONE HUNDRED U.S. DOLLARS (US$100).
        
        --------------------------------------------------------------------------------
        11. INDEMNIFICATION
        --------------------------------------------------------------------------------
        
        11.1 Licensee shall defend, indemnify, and hold harmless Licensor and its
             officers, directors, employees, and agents from and against any and all
             claims, damages, liabilities, costs, and expenses (including reasonable
             attorneys' fees) arising out of or relating to Licensee's use of the
             Software, any use of the Software by any AI Agent acting on Licensee's
             behalf or via Licensee's credentials, or any breach of this Agreement.
        
        --------------------------------------------------------------------------------
        12. EXPORT, SANCTIONS, AND COMPLIANCE
        --------------------------------------------------------------------------------
        
        12.1 Licensee shall comply with all applicable export control, sanctions, and
             anti-corruption laws and shall not export, re-export, or transfer the
             Software to any prohibited destination, entity, or person, or use it for
             any prohibited end use.
        
        --------------------------------------------------------------------------------
        13. GENERAL
        --------------------------------------------------------------------------------
        
        13.1 Governing Law; Venue. This Agreement is governed by the laws of the
             Commonwealth of Pennsylvania, excluding its conflict-of-laws rules. The
             parties submit to the exclusive jurisdiction of the state and federal
             courts located in Pennsylvania.
        
        13.2 Equitable Relief. Licensee acknowledges that any breach of Sections 2, 3,
             4, or 6 would cause irreparable harm for which monetary damages are
             inadequate, and that Licensor is entitled to injunctive relief without the
             need to post bond.
        
        13.3 Assignment. Licensee may not assign or transfer this Agreement or any
             rights or obligations, by operation of law or otherwise, without
             Licensor's prior written consent. Any attempted assignment in violation is
             void. Licensor may freely assign.
        
        13.4 No Waiver. No failure or delay by Licensor in exercising any right waives
             it. Any waiver must be in a writing signed by Licensor.
        
        13.5 Severability. If any provision is held unenforceable, it shall be modified
             to the minimum extent necessary, and the remaining provisions remain in
             full force.
        
        13.6 Entire Agreement. Subject to Section 13.8, this Agreement, together
             with any written authorization issued by Licensor, is the entire agreement
             between the parties regarding the Software and supersedes all prior or
             contemporaneous understandings. Any conflicting or additional terms
             proposed by Licensee are rejected.
        
        13.7 U.S. Government Rights. If Licensee is a U.S. Government entity, the
             Software is "commercial computer software" and "commercial computer
             software documentation," and any use, duplication, or disclosure is
             subject to the restrictions of this Agreement to the extent permitted by
             applicable Federal Acquisition Regulation and agency supplements.
        
        13.8 Amendments; Updated Versions.
        
             (a) Licensor may issue an amended or replacement version of this Agreement
             (an "Updated Agreement"). To be effective as to Licensee, an Updated
             Agreement must: (i) conspicuously identify itself as an updated version of
             this Agreement and state its version number and effective date; (ii) be
             provided to Licensee in full by a method reasonably calculated to give
             Licensee notice; and (iii) be affirmatively accepted by Licensee through
             a click-through acceptance, electronic signature, or other written
             acceptance by a person authorized to bind Licensee. Licensor may condition
             any renewal, extension, update, upgrade, or continued access to the
             Software after the end of the then-current Term upon such acceptance.
        
             (b) Upon Licensee's affirmative acceptance, the Updated Agreement
             supersedes and replaces all prior versions of this Agreement between
             Licensor and Licensee with respect to the Software, effective on the later
             of the Updated Agreement's stated effective date or Licensee's acceptance
             date. Unless the Updated Agreement expressly states otherwise, any then-
             effective written authorization issued by Licensor remains in force only to
             the extent it is consistent with the Updated Agreement.
        
             (c) Until an Updated Agreement becomes effective under this Section 13.8,
             the version of this Agreement previously accepted by Licensee remains
             controlling, subject to Licensor's termination and suspension rights under
             Section 7. Licensor shall retain a reasonably accessible record of each
             version of this Agreement and the date on which Licensee accepted it.
        
        ================================================================================
        For licensing inquiries, contact: legal@dataerai.com
        Dataerai, Inc. — Center Valley, Pennsylvania, USA
        ================================================================================
        
Project-URL: Homepage, https://dataerai.com
Project-URL: Documentation, https://docs.dataerai.com
Project-URL: Source, https://github.com/dataerai/dataerai-toolkit
Keywords: dataerai,data-transfer,sdk,research-data
Classifier: License :: Other/Proprietary License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: dev
Requires-Dist: pytest>=8; extra == "dev"
Requires-Dist: pytest-mock>=3; extra == "dev"
Requires-Dist: ipython>=8; extra == "dev"
Provides-Extra: rest
Requires-Dist: httpx>=0.27; extra == "rest"
Requires-Dist: keyring>=25; extra == "rest"
Provides-Extra: envelope
Requires-Dist: cryptography>=42; extra == "envelope"
Requires-Dist: blake3>=0.3; extra == "envelope"
Provides-Extra: watermark
Requires-Dist: cryptography>=42; extra == "watermark"
Provides-Extra: notebook
Requires-Dist: ipython>=8; extra == "notebook"
Provides-Extra: nn-tensorflow
Requires-Dist: ipython>=8; extra == "nn-tensorflow"
Requires-Dist: tensorflow<2.19,>=2.18; extra == "nn-tensorflow"
Requires-Dist: numpy<3,>=2; extra == "nn-tensorflow"
Requires-Dist: psutil>=5.9; extra == "nn-tensorflow"
Requires-Dist: nvidia-ml-py>=12; extra == "nn-tensorflow"
Provides-Extra: nn-pytorch
Requires-Dist: ipython>=8; extra == "nn-pytorch"
Requires-Dist: torch>=2.6; extra == "nn-pytorch"
Requires-Dist: numpy>=1.26; extra == "nn-pytorch"
Requires-Dist: psutil>=5.9; extra == "nn-pytorch"
Requires-Dist: nvidia-ml-py>=12; extra == "nn-pytorch"
Provides-Extra: bambu
Requires-Dist: imageio-ffmpeg>=0.6; extra == "bambu"
Provides-Extra: ml-io
Requires-Dist: zarr>=3.2; extra == "ml-io"
Requires-Dist: obstore>=0.5; extra == "ml-io"
Requires-Dist: numpy>=1.26; extra == "ml-io"
Requires-Dist: requests>=2.31; extra == "ml-io"
Requires-Dist: s3fs>=2024.0; extra == "ml-io"
Provides-Extra: ml
Requires-Dist: dataerai-sdk[ml-io]; extra == "ml"
Requires-Dist: torch>=2.4; extra == "ml"
Requires-Dist: torchdata>=0.9; extra == "ml"
Requires-Dist: safetensors>=0.4; extra == "ml"
Provides-Extra: croissant
Requires-Dist: mlcroissant>=1.0; extra == "croissant"
Provides-Extra: ml-gpu-cu12
Requires-Dist: dataerai-sdk[ml]; extra == "ml-gpu-cu12"
Requires-Dist: kvikio-cu12>=25.10; extra == "ml-gpu-cu12"
Requires-Dist: cupy-cuda12x>=13; extra == "ml-gpu-cu12"
Provides-Extra: ml-gpu
Requires-Dist: dataerai-sdk[ml-gpu-cu12]; extra == "ml-gpu"
Provides-Extra: serve
Requires-Dist: dataerai-sdk[ml-io]; extra == "serve"
Requires-Dist: onnxruntime>=1.17; extra == "serve"
Requires-Dist: pillow>=10; extra == "serve"
Requires-Dist: safetensors>=0.4; extra == "serve"
Provides-Extra: serve-gpu
Requires-Dist: dataerai-sdk[ml-io]; extra == "serve-gpu"
Requires-Dist: onnxruntime-gpu>=1.17; extra == "serve-gpu"
Requires-Dist: pillow>=10; extra == "serve-gpu"
Requires-Dist: safetensors>=0.4; extra == "serve-gpu"
Provides-Extra: ml-keras
Requires-Dist: dataerai-sdk[ml-io]; extra == "ml-keras"
Requires-Dist: numpy>=2; extra == "ml-keras"
Requires-Dist: keras>=3; extra == "ml-keras"
Requires-Dist: pillow>=10; extra == "ml-keras"
Provides-Extra: ml-hls4ml
Requires-Dist: dataerai-sdk[ml-keras]; extra == "ml-hls4ml"
Requires-Dist: hls4ml>=1.3; extra == "ml-hls4ml"
Provides-Extra: ml-stream
Requires-Dist: numpy>=2; extra == "ml-stream"
Requires-Dist: zarr>=3; extra == "ml-stream"
Provides-Extra: ml-stream-video
Requires-Dist: dataerai-sdk[ml-stream]; extra == "ml-stream-video"
Requires-Dist: av>=12; extra == "ml-stream-video"
Provides-Extra: ml-stream-train
Requires-Dist: dataerai-sdk[ml-stream]; extra == "ml-stream-train"
Requires-Dist: torch; extra == "ml-stream-train"
Provides-Extra: runner
Requires-Dist: requests>=2; extra == "runner"
Requires-Dist: nvidia-ml-py>=12.535; extra == "runner"
Provides-Extra: metaextract
Requires-Dist: numpy>=1.23; extra == "metaextract"
Requires-Dist: h5py>=3.7; extra == "metaextract"
Requires-Dist: igor2>=0.5; extra == "metaextract"
Requires-Dist: lxml>=4.9; extra == "metaextract"
Requires-Dist: hyperspy>=2.0; extra == "metaextract"
Provides-Extra: metaextract-converters
Requires-Dist: dataerai-sdk[metaextract]; extra == "metaextract-converters"
Requires-Dist: xarray>=0.20; extra == "metaextract-converters"
Requires-Dist: netCDF4>=1.6; extra == "metaextract-converters"
Provides-Extra: metaextract-viz
Requires-Dist: dataerai-sdk[metaextract]; extra == "metaextract-viz"
Requires-Dist: plotly>=5; extra == "metaextract-viz"
Requires-Dist: nbformat>=4.2; extra == "metaextract-viz"
Requires-Dist: matplotlib>=3.5; extra == "metaextract-viz"
Provides-Extra: metaextract-checksums
Requires-Dist: dataerai-sdk[metaextract]; extra == "metaextract-checksums"
Requires-Dist: blake3>=0.3; extra == "metaextract-checksums"
Provides-Extra: metaextract-utils
Requires-Dist: dataerai-sdk[metaextract]; extra == "metaextract-utils"
Requires-Dist: requests>=2.25; extra == "metaextract-utils"
Provides-Extra: metaextract-containers
Requires-Dist: dataerai-sdk[metaextract]; extra == "metaextract-containers"
Requires-Dist: netCDF4>=1.6; extra == "metaextract-containers"
Requires-Dist: scipy>=1.9; extra == "metaextract-containers"
Requires-Dist: pyarrow>=10; extra == "metaextract-containers"
Requires-Dist: tifffile>=2022.5; extra == "metaextract-containers"
Requires-Dist: pillow>=9; extra == "metaextract-containers"
Provides-Extra: metaextract-microscopy
Requires-Dist: dataerai-sdk[metaextract]; extra == "metaextract-microscopy"
Requires-Dist: nd2>=0.7; extra == "metaextract-microscopy"
Requires-Dist: czifile>=2019.7; extra == "metaextract-microscopy"
Requires-Dist: readlif>=0.6; extra == "metaextract-microscopy"
Requires-Dist: oiffile>=2021.6; extra == "metaextract-microscopy"
Provides-Extra: metaextract-medical
Requires-Dist: dataerai-sdk[metaextract]; extra == "metaextract-medical"
Requires-Dist: pydicom>=2.3; extra == "metaextract-medical"
Provides-Extra: metaextract-bio
Requires-Dist: dataerai-sdk[metaextract]; extra == "metaextract-bio"
Requires-Dist: pysam>=0.21; extra == "metaextract-bio"
Provides-Extra: metaextract-spectroscopy
Requires-Dist: dataerai-sdk[metaextract]; extra == "metaextract-spectroscopy"
Requires-Dist: brukeropus>=1.1; extra == "metaextract-spectroscopy"
Provides-Extra: metaextract-all
Requires-Dist: dataerai-sdk[metaextract-checksums,metaextract-converters,metaextract-utils,metaextract-viz]; extra == "metaextract-all"
Requires-Dist: dataerai-sdk[metaextract-containers,metaextract-medical,metaextract-microscopy]; extra == "metaextract-all"
Requires-Dist: dataerai-sdk[metaextract-bio,metaextract-spectroscopy]; extra == "metaextract-all"
Dynamic: license-file

# Dataerai Python SDK

Python clients for the Dataerai transfer daemon and the complete public REST
API. The daemon client supports authenticated uploads, downloads, metadata, and
resumable transfers; the optional REST client exposes every public API path.

## Requirements

- Python ≥ 3.10
- The `dataerai` binary must be installed and on `PATH` (or pass `binary_path` explicitly)
- The user must be logged in via `dataerai auth login` before calling `auth_status()` / `upload()` / `download()`

## Installation

```bash
pip install --pre "dataerai-sdk>=0.2.0b1,<0.3"
# or, from source:
pip install -e sdk/python/
```

Optional ML adapters are installed separately so the base SDK stays small:

```bash
pip install --pre "dataerai-sdk[rest]>=0.2.0b1,<0.3"        # Full public REST API
pip install --pre "dataerai-sdk[ml]>=0.2.0b1,<0.3"          # PyTorch Zarr datasets
pip install --pre "dataerai-sdk[ml-keras]>=0.2.0b1,<0.3"    # Keras batches
pip install --pre "dataerai-sdk[ml-hls4ml]>=0.2.0b1,<0.3"   # HLS4ML conversion
pip install --pre "dataerai-sdk[notebook]>=0.2.0b1,<0.3"    # IPython %dataerai magic
pip install --pre "dataerai-sdk[nn-tensorflow]>=0.2.0b1,<0.3" # TensorFlow provenance
pip install --pre "dataerai-sdk[envelope]>=0.2.0b1,<0.3"    # Encrypted-envelope reader
```

## Full public REST API

The daemon client remains the best path for resumable file transfers. Install
the `rest` extra when a script also needs any public console API endpoint:

```python
from dataerai.rest import RestClient

# Reuses credentials written by `dataerai auth login --device`.
with RestClient() as api:
    response = api.request(
        "GET",
        "/api/projects/",
        params={"page_size": 20},
    )
    response.raise_for_status()
    projects = response.json()
```

`request()` accepts every relative `/api/` path, including endpoints added
after this SDK version was published, and returns the raw `httpx.Response`.
`stream()` applies the same security checks while iterating a large response.
Absolute URLs and redirects are rejected so bearer credentials remain scoped to
the configured Dataerai server. For headless jobs, set `DATAERAI_SERVER` and
`DATAERAI_TOKEN`, or provide the CLI-compatible JSON credential object through
`DATAERAI_CREDENTIALS_JSON`.

## Quick start

```python
from dataerai import DataeraiClient

with DataeraiClient(binary_path="/usr/local/bin/dataerai") as client:
    # Check auth
    status = client.auth_status()
    print(f"Logged in as {status.user_email} ({status.user_id}), token expires {status.expires_at}")
    if status.user_id is None:
        raise RuntimeError("Upgrade the dataerai CLI to use user-owned uploads")

    # Upload a user-owned file. owner_id is a UUID, not an email address.
    result = client.upload(
        "/path/to/data.csv",
        title="My dataset",
        owner_type="user",
        owner_id=status.user_id,
        record_type="dataset",
        on_progress=lambda p: print(f"  {p.percent:.0f}%  {p.rate_mbps:.1f} MB/s"),
    )
    print(f"Uploaded  asset_id={result.asset_id}  content_id={result.content_id}")

    # Download it back
    dl = client.download(result.asset_id, dest_dir="/tmp/downloads")
    for f in dl.files:
        print(f"  {f.local_path}  ({f.size:,} bytes)")

    # Read / update metadata
    meta = client.get_metadata(result.asset_id)
    updated = client.set_metadata(
        result.asset_id,
        title="My dataset v2",
        record_type="analysis",
        tags=["csv", "demo"],
    )

    # Link a derived result to the source asset that produced it
    relationship = client.create_relationship(
        result.asset_id,
        "source-asset-id",
        relationship_type="derived_from",
        analysis_mode="non_destructive",
        qualifiers={"tool": "pycroscopy"},
    )
    related_id = relationship.related_asset["id"] if relationship.related_asset else "source-asset-id"
    print(f"Linked via {relationship.type} to {related_id}")
```

## Notebook magic

After signing in with the Dataerai CLI, a notebook only needs a destination
path. The first component names the project; the notebook magic creates that
project when it is missing, then creates any missing collection components
beneath it:

```python
DESTINATION_COLLECTION_PATH = "Research / Experiments / July"

%load_ext dataerai.magics
%dataerai $DESTINATION_COLLECTION_PATH
```

The magic publishes a `dataerai_session` variable. Its uploads automatically
target the selected project and collection:

```python
dataset = dataerai_session.find_asset(
    "did:dataerai:beta:asset:...",
    title="input.csv",
)
dataerai_session.download(dataset, "source-data")
result = dataerai_session.upload(
    "analysis.csv",
    metadata={"component": "derived-output"},
)
```

Use `%dataerai --as run $DESTINATION_COLLECTION_PATH` to choose a different
session variable. Regular Python code can call
`connect_notebook(DESTINATION_COLLECTION_PATH)` instead.

### Trace a notebook run

Add `--trace` to capture every subsequent cell's source, timestamps, stdout,
stderr, structured Python logging records, rich display outputs, returned
value, and error. Uploads and downloads through the session join the same run
automatically, following the run-centered provenance pattern used by the QICK
integration:

```python
%dataerai --trace --notebook analysis.ipynb --title "July analysis" $DESTINATION_COLLECTION_PATH

# Run analysis cells and upload their products through dataerai_session.
result = dataerai_session.upload(
    "analysis.csv",
    record_type="analysis",
    metadata={"component": "derived-output"},
)

%dataerai --finish
```

`%dataerai --finish` waits until its cell completes, then uploads a JSON
execution-log asset with `record_type="log"`. Every uploaded asset carries a
shared `notebook_run_id` plus `dataerai-notebook-trace` and
`notebook-run:<run-id>` tags. The log is linked to each recorded input and
product with `records_telemetry`, and the published result is available as
`dataerai_trace`. Run identity is merged onto an existing same-title asset, so
re-running a fixed notebook or output filename keeps both run tags searchable.
If a traced cell fails, its source, outputs, error, and traceback are published
immediately; an explicit `%dataerai --finish` is not required for that failed
run.

Tracing is deliberately opt-in because cell source and output can contain
secrets or sensitive research data. Review the notebook before enabling it.
The recorder does not read environment-variable values or scan/upload arbitrary
files; a file becomes a recorded product only when the notebook uploads it
through the traced session. Finish an active trace before starting another one
in the same Python kernel.

## TensorFlow neural-network provenance

`TensorFlowProvenanceTracker` implements the complete preservation contract
ported from
[DataFed_TorchFlow](https://github.com/m3-learning/DataFed_TorchFlow). It saves
native `.keras` checkpoints containing model and optimizer state, writes a
machine-readable manifest, captures architecture and hyperparameters, data and
training-code checksums, outcomes, and full runtime/system properties, and
authors `derived_from` links to datasets, the notebook, and prior checkpoints.
Every checkpoint is a first-class `model` asset, so provenance graphs render it
as a dedicated violet compute node with the Model brain-circuit logo. Final
checkpoints may also publish hashed figures as `analysis` assets linked to the
model with `visualizes` relationships. When the notebook session is tracing,
the tracker automatically reuses its run ID across models, manifests, figures,
and the execution log.

For notebook runs that publish many checkpoints or large analysis artifacts,
`%dataerai --request-timeout 120 ...` raises the daemon request/response limit
without changing the separate upload-transfer timeout.

```python
from dataerai.nn import TensorFlowProvenanceTracker

tracker = TensorFlowProvenanceTracker(
    model,
    session=dataerai_session,
    run_name="mnist-mlp",
    dataset_references=[
        {
            "asset_id": dataset_asset.asset_id,
            "source_asset_did": SOURCE_ASSET_DID,
            "filename": DATA_PATH.name,
            "sha256": dataset_sha,
        }
    ],
    notebook_path=NOTEBOOK_PATH,
    notebook_asset_id=notebook_asset.asset_id,
)
history = model.fit(
    x_train,
    y_train,
    epochs=5,
    callbacks=[tracker.callback(training_parameters={"batch_size": 128})],
)
final = tracker.save_checkpoint(
    epoch=5,
    label="final",
    record_name="mnist-mlp",
    metrics={"test_accuracy": 0.91},
    training_parameters={"batch_size": 128, "epochs": 5},
    outcomes={"sample_predictions": sample_predictions},
    analysis_artifacts=[
        {
            "path": "artifacts/training-curves.png",
            "kind": "training-curves",
            "caption": "Training and validation metrics across five epochs.",
            "media_type": "image/png",
        }
    ],
)
```

## PyTorch neural-network provenance

Install the PyTorch adapter when a training run must preserve native model and
optimizer state:

```bash
pip install "dataerai-sdk[nn-pytorch]"
```

`PyTorchProvenanceTracker` emits the same DataFed_TorchFlow preservation
concepts as the TensorFlow tracker while keeping framework-native checkpoint
validation inside the adapter. Each `.pt` file contains only the versioned
schema, run/epoch identity, model `state_dict`, and optimizer `state_dict`.
Before any upload, the tracker reloads the file on CPU with
`weights_only=True`; a checkpoint that cannot pass that restricted loader is
rejected.

```python
import torch

from dataerai.nn import PyTorchProvenanceTracker

model = MyModel()
optimizer = torch.optim.Adam(model.parameters(), lr=1e-3)
tracker = PyTorchProvenanceTracker(
    model,
    optimizer,
    session=dataerai_session,
    run_name="mnist-cnn",
    dataset_references=[
        {
            "asset_id": dataset_asset.asset_id,
            "source_asset_did": SOURCE_ASSET_DID,
            "filename": DATA_PATH.name,
            "sha256": dataset_sha,
        }
    ],
    notebook_path=NOTEBOOK_PATH,
    notebook_asset_id=notebook_asset.asset_id,
)

for epoch in range(1, 6):
    train_one_epoch(model, optimizer)
    tracker.save_checkpoint(
        epoch=epoch,
        label="epoch",
        metrics={"loss": current_loss},
        training_parameters={"learning_rate": 1e-3, "epochs": 5},
    )
```

With a `NotebookSession`, checkpoints are published as `model` assets.
Manifests and optional analysis figures are ordinary `analysis` assets.
Qualified `derived_from`, `describes`, and `visualizes` relationships connect
them to datasets, the training notebook, and the preceding checkpoint. Without
a session, the identical checkpoint and manifest contracts are written locally.

System capture intentionally does not enumerate environment variables. It
records OS/host, CPU and memory properties, framework-visible devices, optional
NVML GPU memory, Python/framework/package versions, and only a small allowlist
of determinism and thread settings.

## API reference

### Encrypted envelopes

The optional reader keeps OAuth credentials in the Go daemon. Python receives
only the audited per-envelope DEK over the owner-only local IPC endpoint:

```python
from dataerai import envelope

with envelope.open("data.denv") as env:
    print(env.audit)
    env.verify_signature()  # server-bound cache; offline while fresh (5 minutes)
    env.unlock()            # daemon IPC; audited purpose="open"
    print(env.namelist())
    with env.open("member.csv") as member:
        first_kib = member.read(1024)
    env.extractall("out")   # atomic; failures leave no partial output
```

Populate or refresh the public-key cache with
`dataerai envelope verify data.denv`. A new Python process must call `unlock()`
again; neither the SDK nor daemon persists the DEK.

The whole path — parsing, fresh-cache verification, `unlock()` and
`extractall()` — is portable. The daemon transport is an `AF_UNIX` socket on
POSIX and a named pipe on Windows, so no platform needs to fall back to the Go
CLI.

### `DataeraiClient(*, socket_path, binary_path, auto_start, start_timeout_s, request_timeout_s)`

| Parameter | Default | Description |
|---|---|---|
| `socket_path` | First match wins: `$DATAERAI_SOCKET` → on Windows `\\.\pipe\dataerai-transfer-<USERNAME>` → `$XDG_RUNTIME_DIR/dataerai-transfer.sock` when that is set (systemd Linux, usually under `/run/user/<uid>/`) → `<tempdir>/dataerai-transfer.sock` | Daemon IPC endpoint — an `AF_UNIX` socket on POSIX, a named pipe on Windows |
| `binary_path` | `None` | Path to the `dataerai` binary (required for `auto_start`) |
| `auto_start` | `True` | Spawn the daemon if no daemon is listening on the endpoint |
| `start_timeout_s` | `10.0` | Seconds to wait for the daemon endpoint to become available |
| `request_timeout_s` | `30.0` | Per-request timeout in seconds |

Use as a context manager (`with DataeraiClient(...) as client:`) for automatic cleanup, or call `client.connect()` / `client.close()` manually.

### Methods

| Method | Returns | Description |
|---|---|---|
| `auth_status()` | `AuthStatus` | Logged-in user ID, email, and token expiry; `user_id` is `None` with older daemons |
| `list_tree()` | `list[TreeNode]` | List visible workspaces, projects, and collections |
| `list_projects()` | `list[Project]` | Projects from the tree (see the write-access caveat below) |
| `list_collections(*, owner_type=None, owner_id=None)` | `list[Collection]` | Collections from the tree, flattened and filterable |
| `list_collection_assets(collection_id, ...)` | `AssetSearchPage` | One page of a collection's assets |
| `get_collection_manifest(collection_id)` | `CollectionManifest` | Every **file** in a collection, with which ones a sync will not deliver |
| `find_collection_assets(collection_id, ...)` | `list[AssetSummary]` | Every asset in a collection, paging until exhausted |
| `create_collection(title, *, owner_type, owner_id, parent_id=None)` | `Collection` | Create a collection |
| `ensure_collection_path(path, *, create_project=False, project_description="")` | `CollectionDestination` | Resolve `Project / Collection / ...`, creating missing collection components and optionally the project |
| `search_assets(query, ...)` | `AssetSearchPage` | Search one page of visible assets |
| `find_assets(query, ...)` | `list[AssetSummary]` | Search and collect every result page |
| `create_project(name, *, description="")` | `Project` | Create a project (+ root collection, owner membership, default allocation); requires write scope |
| `list_allocations(owner_id, *, owner_type="project")` | `list[Allocation]` | A project's storage allocations, with quota headroom |
| `upload(local_path, *, title, owner_type, owner_id, record_type=None, ...)` | `UploadResult` | Upload a file with an optional server-validated record type; blocks until complete |
| `upload(local_path, *, title, owner_type, owner_id, record_type=None, ...)` | `UploadResult` | Upload one file — or a sequence of paths as **one multi-file asset**; blocks until complete |
| `upload_many(items, *, owner_type, owner_id, concurrency=4, ...)` | `BulkUploadResult` | Upload many assets concurrently, each with its own metadata; reports partial failure |
| `update_content(asset_id, local_path, ...)` | `UploadResult` | Attach a **new content version** to an existing asset, by ID; blocks until complete |
| `list_content_versions(asset_id)` | `list[ContentVersion]` | List an asset's content versions (in-flight ones first, then available newest-first) |
| `download(asset_id, dest_dir, ...)` | `DownloadResult` | Download latest asset content; blocks until complete |
| `download_collection(collection_id, dest_dir, ...)` | `CollectionDownloadResult` | Download every asset in a collection recursively, preserving sub-collections as sub-directories; blocks until all transfers complete |
| `list_transfers()` | `list[TransferSummary]` | Transfers the local daemon is tracking |
| `list_server_transfers(...)` | `ServerTransferPage` | Transfers the console records, including ones started elsewhere |
| `cancel_transfer(transfer_id)` | `str` | Cancel a transfer; waits for the daemon to confirm |
| `pause_transfer(transfer_id)` | `None` | Pause a transfer (**fire-and-forget**) |
| `resume_transfer(transfer_id, credentials=None)` | `None` | Resume a paused transfer (**fire-and-forget**) |
| `shutdown_daemon()` | `None` | Stop the daemon **for every client on the machine**, then close this client (**fire-and-forget**) |
| `get_metadata(asset_id)` | `AssetMetadata` | Retrieve asset metadata |
| `set_metadata(asset_id, **fields)` | `AssetMetadata` | Update metadata fields, including `record_type` |
| `set_metadata_many(updates, *, concurrency=4)` | `BulkMetadataResult` | Apply a per-asset metadata edit to many assets; reports partial failure |
| `create_relationship(from_asset_id, to_asset_id, rel_type=None, *, relationship_type=None, ...)` | `Relationship` | Create a directed provenance link between two assets |
| `on(event, handler=None)` | handler / decorator | Subscribe to a daemon event, or to `"*"` for all (see [Daemon events](#daemon-events)) |
| `off(event, handler)` | `bool` | Unsubscribe; `False` if it was not registered |
| `sync_collection(collection_id, dest_dir, *, prune=False, dry_run=False, ...)` | `SyncResult` | Incrementally sync a collection into a local folder; **`prune` deletes local files** |
| `delete_relationship(from_asset_id, relationship_id)` | `None` | Delete an edge from its **source** asset — the inverse of `create_relationship()` |

### Choosing an allocation

`upload()` and `update_content()` accept an `allocation_id`; this is how you get
one:

```python
for a in client.list_allocations(project_id):
    print(a.allocation_id, a.alias, a.repository_name,
          f"{a.bytes_free:,} B free", f"{a.records_free:,} records free",
          "FULL" if a.is_full else "")
```

Checking before you upload is worth the round trip. A full allocation doesn't
fail at the point of cause — uploads are rejected, and a rejected upload can
leave a contentless asset that surfaces much later as "no downloadable
content".

> **`is_full` is necessary, not sufficient.** `True` means an upload will be
> rejected. `False` does *not* guarantee one succeeds: the console also
> hard-blocks on billing suspension, enforces a temporary grace ceiling instead
> of `data_volume_limit` during a grace window, and walks up parent
> allocations — none of which the daemon reports, so none are visible here.
> Read `False` as "the reported quota has room".

A limit of `0` means zero capacity, not unlimited — the console's gate is a
plain `current + incoming > limit` with no special case for zero. A missing limit
likewise reads as full rather than unlimited: fail-closed is the safe
direction, because wrongly reporting room is the failure this is meant to
prevent.

> **An empty list does not mean uploads will fail.** The server resolves an
> upload's allocation through a cascade — the owner's allocation on its
> preferred repository, else the owner's default, else, for a project-owned
> asset, one held by the project's *owning user*. That last step is invisible
> here, because the daemon lists only what the project holds directly. So an
> upload with no `allocation_id` may land in an allocation that is not in this
> list; passing one from here pins the destination instead.
### Listing what you have

`list_collection_assets()` / `find_collection_assets()` enumerate a collection's
contents without downloading them — the alternative before they existed was
`download_collection()`, which answers the question by fetching everything:

```python
for asset in client.find_collection_assets(collection_id):
    print(asset.asset_id, asset.title, asset.size_bytes, asset.has_content)
```

`find_collection_assets()` pages until exhausted, and raises rather than looping
if the server claims another page without a cursor or repeats one.

> **`list_projects()` and `list_collections()` show what you can upload into,
> not everything you can read.** Both are projections of `list_tree()`, and the
> daemon omits any project you lack `WRITE_METADATA` on — so a project you can
> read but not write to is absent, and so are its collections. There is no
> SDK call that lists read-only projects today.

`list_projects()` populates only `project_id`, `name` and `root_collection_id`;
`tree.list` carries nothing else, so the remaining `Project` fields are `None`.
Collections shared with you directly are attached to the personal node, so one
returned under `owner_type="user"` may carry another project's
`owner_project_id`.

For headless notebooks, pass `record_type` to `upload()` or `set_metadata()`.
Both operations use the authenticated transfer daemon, so integrations do not
need to read the CLI credential store or expose access tokens to Python.

### What a collection actually contains

`list_collection_assets()` enumerates *assets*. `get_collection_manifest()`
enumerates the **files** inside them, flattened, with each one's position in the
collection tree — and, crucially, which of them a download or sync will
**silently not deliver**:

```python
manifest = client.get_collection_manifest(collection_id)
print(f"{manifest.total_files} files, {manifest.total_bytes:,} bytes")

for entry in manifest.undownloadable:
    why = "external" if entry.external else "no checksum on the server"
    print(f"  will NOT arrive: {entry.relative_path}  ({why})")
```

This is the answer to *"why does my synced folder have fewer files than the
collection?"*. A sync omits two kinds of entry and reports neither per-file:
**external** entries, whose bytes live outside Dataerai, and entries the server
holds **no checksum** for, which the daemon refuses to write because it cannot
verify them (the `skipped_unverified` count). The daemon collapses both into one
`is_downloadable` flag, so `external` is what tells them apart.

The manifest is **complete or it raises** — the daemon assembles every page
itself and rejects a result whose totals or snapshot shifted while paging, which
is what makes it safe to use as the expectation you check a sync against. Note
it carries **no checksums**; the checksum is consumed to compute
`is_downloadable`, not forwarded.

`collection_id` must be a **canonical** UUID here — lower-case, hyphenated, no
braces or `urn:` prefix. The daemon rejects other spellings rather than
normalising them.

`download_collection()` raises on a failed or timed-out transfer, but an asset
the daemon could never queue is *reported* rather than raised: its ID lands in
`result.failed_assets`, and it is absent from both `result.transfers` and
`result.asset_count`. Check that field before treating the download as complete.

```python
result = client.download_collection(collection_id, "/data/out")
if result.failed_assets:
    raise RuntimeError(f"{len(result.failed_assets)} assets did not download: "
                       f"{result.failed_assets}")
```

### Watching and controlling transfers

The daemon runs a real job engine behind `upload()` and `download()`: chunked
multipart, parts in parallel, resume-from-partial across restarts, and at most
four transfers at once with the rest queued. These expose it:

```python
for t in client.list_transfers():
    print(t.transfer_id, t.status, t.direction, f"{t.percent:.0f}%")

# Transfers started anywhere — the web app, another device, a runner
page = client.list_server_transfers(type="upload")
print(page.count, "total;", [t.executor for t in page.transfers])
```

> **`pause_transfer()` and `resume_transfer()` are fire-and-forget, and
> therefore silent on failure.** The daemon writes nothing on success, so they
> cannot wait for an acknowledgement — and one naming an unknown transfer is
> discarded without raising. This mirrors the Node SDK rather than inventing a
> Python-only contract. Confirmation is asynchronous: the worker emits
> `transfer.paused`, and `list_transfers()` then shows `paused`.
>
> `cancel_transfer()` is different — the daemon acks it, so it waits and raises
> `ERR_TRANSFER_NOT_FOUND` if the transfer is unknown.

### Stopping the daemon

`shutdown_daemon()` completes the set of fire-and-forget commands the Node SDK
sends. The daemon logs the request, signals its own shutdown and closes the
connection without replying.

> ⚠️ **The daemon is shared machine-wide.** Stopping it aborts in-flight
> transfers for *every* client — the desktop app and any running CLI included.
> Use `close()` unless you specifically mean to stop the daemon, e.g. tearing
> down an ephemeral environment or a test fixture.

The client is closed afterwards: one left open against a stopped daemon fails
every later call with a connection error that says nothing about the cause.
Build a new client to reconnect — one with `auto_start` and a `binary_path`
respawns the daemon.

**What a blocking call sees.** Cancelling a transfer that an `upload()` or
`download()` is waiting on makes *that* call raise `DaemonError` with
`code == "cancelled"` — the daemon emits `transfer.failed` with that code
rather than a `transfer.cancelled` message, so check the code rather than
expecting a clean return. **Pausing is different and easier to get wrong:** a
paused transfer sends no terminal event, so a blocked `upload()` keeps waiting
and eventually raises `DaemonTimeoutError` after its `transfer_timeout_s`.
Pausing does not extend that budget.

Cancelling an upload that is still running needs its transfer id, and
`upload()` does not return until it finishes — so capture the id from a
progress event:

```python
seen = {}
client.upload(path, ..., on_progress=lambda e: seen.setdefault("id", e.transfer_id))
# ... from another thread, once seen has an id:
client.cancel_transfer(seen["id"])
```

### Keeping a local folder in sync

`download_collection()` fetches everything, every time. `sync_collection()`
compares the collection against what is already on disk and moves only what
changed, keeping state under the destination so it works across runs:

```python
result = client.sync_collection(
    collection_id,
    "/data/my-collection",
    on_progress=lambda p: print(f"  {p.percent:.0f}%  {p.path}"),
)
print(f"{result.downloaded} downloaded, {result.pruned} removed, "
      f"{result.elapsed_ms} ms")

if not result.is_complete:
    print(
        f"{result.failed} files could not be fetched; "
        f"{result.conflicts} local conflicts were preserved"
    )
if result.skipped_unverified:
    print(f"{result.skipped_unverified} withheld — the server had no checksum")
```

> ⚠️ **`prune=True` deletes local files.** Anything under the destination the
> collection no longer contains is removed. Preview it first — `dry_run=True`
> computes the same plan, writes nothing, deletes nothing, and does not even
> create a missing destination:
>
> ```python
> preview = client.sync_collection(cid, dest, prune=True, dry_run=True)
> print(f"{preview.plan.prune_candidates} local files would be deleted")
> print(f"{preview.plan.to_download} would be fetched ({preview.plan.bytes:,} bytes)")
> ```

**A partial sync does not raise.** Files that could not be fetched are counted
in `failed`. Locally modified files that the daemon preserved are counted in
`conflicts`. Either makes `is_complete` false because the destination was not
brought fully in line. Files that were never going to arrive are counted
separately in `skipped_external` and `skipped_unverified`, and deliberately do
not affect `is_complete`, since the plan declared them up front.
`get_collection_manifest()` names them individually.

**A sync cannot be cancelled.** The daemon runs it detached from the request, so
`sync_timeout_s` bounds how long you *wait* — not the sync, which keeps running
and keeps writing. Closing the client also abandons only your wait if the shared
daemon remains alive. Only one sync may be active per destination; a second
raises `DaemonError`.

### Updating an existing asset

Metadata and content are updated by two different calls, both taking an
`asset_id`:

```python
client.set_metadata(asset_id, tags=["v2"])              # metadata
client.update_content(asset_id, "v2.csv")               # a new content version
```

Earlier versions are kept. List them and fetch an older one by ID:

```python
for v in client.list_content_versions(asset_id):
    print(v.content_id, v.status, v.size_bytes, v.created_at)

client.download(asset_id, "out/", content_id="<older content_id>")
```

A version appears in that list as soon as its upload starts, so an in-flight
one shows up with `status="uploading"` and `is_downloadable == False` — it has
no bytes behind it yet. They are surfaced rather than hidden so a concurrent
upload doesn't look like nothing is happening.

**Order:** in-flight versions come first, then the available ones newest-first.
So `[0]` is not reliably the newest downloadable version while an upload is
running — filter rather than index:

```python
newest = next(v for v in client.list_content_versions(asset_id)
              if v.is_downloadable)
```

> **`upload()` with a title that already exists updates the content but
> silently drops the metadata.** The daemon upserts by
> `(collection, title, owner)`; when that matches, the console returns the
> existing asset *unchanged* and attaches a new content version. The bytes
> land, but `description`, `alias`, `record_type`, `tags` and `metadata` from
> that call are ignored with no error. Use `update_content()` to add a version
> to a known asset, and `set_metadata()` to change fields — both take an
> `asset_id`, so neither can hit the wrong asset.

### Bulk operations

`upload()` takes a sequence of paths to build **one asset from several files** —
the daemon puts them in a single `AssetContent` and moves them under one
transfer. Filenames must be unique within the asset, since the filename is the
server-side object key:

```python
client.upload(["run.csv", "run.json"], title="Run 7",
              owner_type="project", owner_id=project_id)
```

`upload_many()` is the other axis: **one asset per item**, uploaded with bounded
concurrency. The keyword arguments are batch defaults and every `UploadItem`
field overrides the default of the same name, so assets that differ in only a
field or two stay readable:

```python
from dataerai import UploadItem

result = client.upload_many(
    [
        UploadItem("s1.csv", title="Sample 1", metadata={"well": "A1"}),
        UploadItem("s2.csv", title="Sample 2", metadata={"well": "B2"}),
        UploadItem(["s3.csv", "s3.json"], title="Sample 3"),
    ],
    owner_type="project",
    owner_id=project_id,
    record_type="dataset",            # applies to all three
    metadata={"run": "2026-07-28"},   # merged into each item's own metadata
    concurrency=4,
)
result.raise_for_failures()
```

`metadata` is **shallow-merged** — batch keys first, then the item's, item wins
on collision. Every other field (including `tags`) replaces the default
outright, so a shared tag can be dropped for one item.

`set_metadata_many()` does the same for metadata-only edits:

```python
from dataerai import MetadataUpdate

client.set_metadata_many([
    MetadataUpdate(a_id, tags=["qc-pass"], metadata={"well": "A1"}),
    MetadataUpdate(b_id, tags=["qc-fail"], metadata={"well": "B2"}),
]).raise_for_failures()
```

Note this is **client-side fan-out, not a protocol batch**: the daemon exposes
only a single-asset `asset.metadata.set`, so it issues one request per asset.
You get bounded concurrency, one call, and one failure report — not fewer
round-trips.

> **Both bulk calls finish the batch instead of aborting on the first failure.**
> Aborting halfway through a large batch leaves you worse off than finishing and
> reporting, so failures land in `result.failed` (each with the item's input
> `index`) and successes in `result.succeeded`. A caller that ignores `failed`
> will read a partial batch as a complete one — check `result.ok` or call
> `result.raise_for_failures()`, which raises `BulkOperationError`.

`result.results` is positionally aligned with the input list — `results[i]` is
the outcome of item `i`, or `None` if it failed. That is how you map an asset
back to the item that produced it, which matters precisely because each item
carries its own metadata:

```python
for item, uploaded in zip(items, result.results):
    if uploaded is not None:
        print(item.title, "->", uploaded.asset_id)
```

`succeeded` and `failed` are convenience views over the same outcomes, both in
input order rather than completion order.

**On `concurrency`:** raising it past the default of 4 does not buy throughput.
The daemon runs at most 4 transfers at once and queues the rest
(`defaultMaxConcurrent` in `cli/internal/transfer/manager.go`, with no
configuration knob), so a higher value only parks more SDK worker threads on
transfers the daemon has not started. Note also that `transfer_timeout_s` starts
when the daemon *accepts* a transfer, not when it starts moving bytes — time
spent queued counts against it, so a large batch with a lowered timeout can fail
items purely from queueing.

### ML adapters

The `dataerai.ml` package reads Dataerai-hosted Zarr image stores directly from
the short-lived `/zarr-access/` bundle. PyTorch users can keep using
`ZarrAssetDataset` / `ZarrOriginalsDataset`; Keras users get matching batch
datasets:

```python
from dataerai.ml import access
from dataerai.ml.keras_dataset import ZarrOriginalsSequence

creds = access.load_credentials()
train = ZarrOriginalsSequence(
    creds,
    asset_id="dataset-asset-id",
    size=64,
    batch_size=32,
    shuffle=True,
    label_mode="categorical",
)

model.fit(train, epochs=5)
```

For HLS4ML conversion, keep inputs in Keras' channels-last layout:

```python
from dataerai.ml.hls4ml import convert_keras_model

hls_model = convert_keras_model(
    model,
    output_dir="hls-project",
    project_name="dataerai_model",
    data=train,              # writes input/output .npy testbench arrays
    testbench_batches=2,
)
```

### `create_relationship()` arguments

Authors a directed provenance edge `from_asset_id` → `to_asset_id`. You need
write access to the source and read access to the target. `rel_type` is a
free-form verb describing the source's role, e.g. `"analysis_of"` or
`"acquired_with"`. Notebook integrations can pass the same value with the
keyword-only alias `relationship_type`.

| Argument | Type | Description |
|---|---|---|
| `from_asset_id` | `str` | Source (dependent) asset — the edge starts here (required) |
| `to_asset_id` | `str` | Target (origin) asset — the edge points here (required) |
| `rel_type` | `str \| None` | Free-form relationship type, ≤255 chars (required unless `relationship_type` is provided) |
| `relationship_type` | `str \| None` | Keyword-only alias for `rel_type` |
| `analysis_mode` | `str \| None` | `non_destructive`, `altering`, `destructive`, `in_situ`, `ex_situ`, `invasive`, `non_invasive` |
| `qualifier_note` | `str \| None` | Free-text note on the relationship |
| `qualifier_time` | `str \| None` | ISO-8601 timestamp |
| `qualifiers` | `dict \| None` | JSON-serializable extra qualifiers |

```python
# Link a processed result back to the raw data it came from.
client.create_relationship(analysis.asset_id, raw.asset_id, "analysis_of",
                           analysis_mode="non_destructive")
```

### `upload()` keyword arguments

| Argument | Type | Description |
|---|---|---|
| `title` | `str` | Asset title (required) |
| `owner_type` | `str` | `"project"` or `"user"` (required) |
| `owner_id` | `str` | Owner entity ID (required) |
| `description` | `str \| None` | Free-text description |
| `alias` | `str \| None` | Short identifier |
| `tags` | `list[str] \| None` | Tag list |
| `metadata` | `dict \| None` | Arbitrary key-value metadata |
| `collection_id` | `str \| None` | Collection to add the asset to |
| `chunk_size_mb` | `int \| None` | Override default 64 MiB chunk size |
| `on_progress` | `Callable[[ProgressEvent], None] \| None` | Progress callback |
| `transfer_timeout_s` | `float` | Max seconds to wait for completion (default 3600) |

### Progress events

```python
@dataclass
class ProgressEvent:
    transfer_id: str
    bytes_done: int
    bytes_total: int
    chunk_index: int
    chunk_count: int
    rate_bps: float
    file_index: int
    file_name: str

    @property
    def percent(self) -> float: ...   # 0–100

    @property
    def rate_mbps(self) -> float: ...
```

### Daemon events

`on_progress=` reports on one call's own transfer. `client.on()` subscribes to
the daemon itself, which is the only way to observe a collection sync, a
non-fatal `transfer.error`, or a pause.

```python
client = DataeraiClient()
client.connect()

@client.on("transfer.error")
def on_error(evt):
    # Not a failure: when evt.retrying is set the daemon retries by itself.
    print(f"{evt.transfer_id}: {evt.message} (retrying={evt.retrying})")

@client.on("collection.sync_complete")
def on_sync(evt):
    if evt.skipped_unverified:
        print(f"{evt.skipped_unverified} files skipped — no checksum offered")
```

The daemon **broadcasts to every connected client**, so handlers also see work
started by the desktop app or the CLI, not just this process.

> **Handlers run on the reader thread and must return promptly.** The daemon
> buffers 64 events per connection and drops the overflow rather than blocking
> other clients — a slow handler loses events outright, and they are not
> redelivered. Queue the work; don't do it in the handler.
>
> In particular, **do not call a blocking client method from a handler.** The
> reader thread is what delivers responses, so such a call waits for itself and
> stalls until its request timeout. Put the id on a queue and act on it from
> your own thread.

An unknown event name raises `ValueError` at registration rather than never
firing. `dataerai.EVENT_NAMES` holds the full set.

Subscribe to `"*"` to see everything — for logging, or to find out what the
daemon actually emits during an operation. A wildcard handler gets the same
parsed event object, so use `event_name()` when the name matters:

```python
from dataerai import event_name

client.on("*", lambda e: print(event_name(e), e))
```

Handlers registered for an event's own name run before wildcard handlers.

Exceptions raised inside a handler never reach the reader thread, but they are
**logged** to the `dataerai.events` logger rather than discarded — a handler
with the wrong signature raises on every event, and a silent version of that is
indistinguishable from the daemon never sending one.

Each distinct failure is reported **once**, with its traceback. A broken handler
fails on every matching event and `transfer.progress` arrives continuously, so
reporting every occurrence would bury the first traceback under thousands of
copies. Two different broken handlers still get one report each. Silence the
lot with `logging.getLogger("dataerai.events").setLevel(logging.CRITICAL)`.

> **Watching an `upload()` or `download()` finish? Subscribe to
> `asset.upload_complete` / `asset.download_complete`, not `transfer.complete`.**
> For an asset transfer the daemon sends the `asset.*` event *instead of* the
> generic one — never both — so `transfer.complete` stays silent for exactly the
> transfers the SDK starts. The same substitution applies to `transfer.failed`.

| Event | Payload | Notes |
|---|---|---|
| `transfer.progress` | `ProgressEvent` | Also available per-call via `on_progress=` |
| `transfer.file_complete` | `FileCompleteEvent` | One file of a multi-file transfer |
| `transfer.complete` | `TransferCompleteEvent` | **Not sent for asset transfers** — see the note above |
| `transfer.error` | `TransferErrorEvent` | **Non-fatal.** `retrying` means the daemon retries by itself |
| `transfer.failed` | `TransferFailedEvent` | Terminal, and **not sent for asset transfers**. A cancel arrives here too — there is no `transfer.cancelled` event |
| `transfer.paused` | `TransferPausedEvent` | The confirmation for a fire-and-forget pause |
| `asset.upload_complete` | `UploadCompleteEvent` | |
| `asset.download_complete` | `DownloadCompleteEvent` | |
| `asset.upload_failed` | `UploadFailedEvent` | |
| `asset.download_failed` | `DownloadFailedEvent` | |
| `collection.sync_progress` | `CollectionSyncProgressEvent` | Keyed by `sync_id`, not `transfer_id` |
| `collection.sync_complete` | `CollectionSyncCompleteEvent` | `skipped_unverified > 0` means files were **not** written |

Names are the daemon's own wire names, so the string here is the string in the
protocol reference and in the daemon's logs. Porting from the Node SDK, which
spells them `transfer:fileComplete`: drop the camelCase and use `.` for `:`.

### Error types

| Exception | When raised |
|---|---|
| `DaemonError(code, message)` | Daemon returned a coded error (see `code` attribute) |
| `DaemonTimeoutError` | Request or transfer exceeded the configured timeout |
| `ConnectionError` | Daemon disconnected unexpectedly |
