Metadata-Version: 2.5
Name: signal-dataset
Version: 0.4.0
Summary: An open, profile-free format and Python library for large-scale multidimensional signal datasets.
Project-URL: Documentation, https://sds.docs.superpose.us
Project-URL: Changelog, https://github.com/superpose-labs/signal-dataset/blob/main/CHANGELOG.md
Project-URL: Issues, https://github.com/superpose-labs/signal-dataset/issues
Project-URL: Repository, https://github.com/superpose-labs/signal-dataset
Author: signal-dataset maintainers
License-Expression: Apache-2.0
License-File: LICENSE
Keywords: arrayrecord,dataset,file-format,immutable,machine-learning,multidimensional,safetensors,signal-processing
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: MacOS
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: System :: Archiving
Classifier: Typing :: Typed
Requires-Python: <3.14,>=3.11
Requires-Dist: array-record<0.9,>=0.8.3
Requires-Dist: numpy<2.5,>=1.26
Requires-Dist: safetensors<0.9,>=0.5
Provides-Extra: gcs
Requires-Dist: google-cloud-storage<4,>=3; extra == 'gcs'
Provides-Extra: s3
Requires-Dist: boto3<2,>=1.35.69; extra == 's3'
Description-Content-Type: text/markdown

# signal-dataset

[![PyPI](https://img.shields.io/pypi/v/signal-dataset.svg)](https://pypi.org/project/signal-dataset/)
[![Python versions](https://img.shields.io/pypi/pyversions/signal-dataset.svg)](https://pypi.org/project/signal-dataset/)
[![License](https://img.shields.io/pypi/l/signal-dataset.svg)](LICENSE)
[![CI](https://github.com/superpose-labs/signal-dataset/actions/workflows/ci.yml/badge.svg)](https://github.com/superpose-labs/signal-dataset/actions/workflows/ci.yml)
[![Docs](https://img.shields.io/badge/docs-sds.docs.superpose.us-blue)](https://sds.docs.superpose.us)

Immutable, indexed storage for multidimensional signal records on local filesystems, Google Cloud
Storage, and Amazon S3. Numerical fields use SafeTensors, and ArrayRecord provides random access.

Use it when you have many arbitrary-rank signal captures, several machines writing them at once,
and readers that must never see a half-written dataset.

> **Status:** experimental 0.x software. The Python API may make documented breaking changes
> between minor releases. Persisted-format compatibility is versioned separately.

```python
import numpy as np
import signal_dataset as sds

record = sds.Record(
    id="capture-0042",
    fields={
        "samples": sds.Field(
            np.zeros((4, 4096), dtype=np.complex64),
            axes=(
                sds.Axis("channel", 4),
                sds.Axis(
                    "time",
                    4096,
                    coordinate=sds.Coordinate(start=0, step=5e-8, unit="s"),
                ),
            ),
        )
    },
    metadata={"sample_rate_hz": 20_000_000},
)

shard = sds.write_shard([record], "captures.sds", work_id="worker-000")
dataset = sds.publish(
    "captures.sds",
    [shard],
    dataset_id="captures",
    snapshot_id="run-001",
)

assert dataset[0].id == "capture-0042"
assert dataset.record_metadata[0]["samples"].shape == (4, 4096)

for descriptor in dataset.iter_record_metadata():
    print(descriptor.id)

for full_record in dataset.iter_records():
    assert full_record["samples"].data.shape == (4, 4096)
```

Reading `dataset[i]` fetches sample tensors. Reading `dataset.record_metadata[i]` does not, which
makes it cheap to plan a shuffle or a filter across a whole dataset before fetching anything.
`iter_records()` and `iter_record_metadata()` stream the same order in bounded batches.

Readers follow generation-pinned manifests and never list a directory or a bucket prefix.

## Install

```bash
pip install signal-dataset
pip install 'signal-dataset[gcs]'   # gs:// roots
pip install 'signal-dataset[s3]'    # s3:// roots
```

The `uv` equivalents are `uv add signal-dataset` and so on.

Python 3.11 through 3.13 on Linux and macOS. Windows is not supported, because atomic publication
uses POSIX filesystem primitives.

## Documentation

Full documentation is at **[sds.docs.superpose.us](https://sds.docs.superpose.us)**.

- [Why signal-dataset](https://sds.docs.superpose.us/why.html), including how it compares to HDF5,
  Zarr, Parquet, WebDataset, and TFRecord.
- [Quickstart](https://sds.docs.superpose.us/quickstart.html)
- [API reference](https://sds.docs.superpose.us/api.html)
- [Format specification](https://sds.docs.superpose.us/specification.html)

## Design

- Immutable, root-last publication: readers observe no dataset or a complete snapshot.
- Lazy ordinal random access without directory or bucket-prefix discovery.
- Tensor-free aligned metadata reads.
- Storage and shard-container extension contracts.
- Profile-free records: no modality, training framework, or split policy is embedded in the core.

The project does not provide mutable datasets, distributed scheduling, authorization, retention,
or batching policy. See the [architecture](docs/architecture.md), [format compatibility
contract](docs/concepts/compatibility.md), and [roadmap](ROADMAP.md).

## Continuous integration

CI runs the test suite with a coverage floor, `ruff`, strict `mypy`, and the installed-wheel,
extras, and sdist journeys on every pull request.

**macOS is not covered by CI**, so running the suite locally is the only signal there. That
matters more than it would for a pure-Python package, because this one is POSIX-specific and the
platforms differ exactly where it works: `/dev/fd` semantics on a staged descriptor, `os.link`,
`fcntl.flock`, directory `fsync`, and whether an unlinked open file keeps a link count.

```bash
uv sync --all-extras --dev --locked
uv run pytest
```

## Releasing

See [RELEASING.md](RELEASING.md).

## Community

Read [CONTRIBUTING.md](CONTRIBUTING.md) before proposing changes. Use GitHub Issues for reproducible
bugs and design discussion, and follow [SECURITY.md](SECURITY.md) for private vulnerability reports.
Participation is governed by the [Code of Conduct](CODE_OF_CONDUCT.md).
