Metadata-Version: 2.5
Name: sensorlint
Version: 0.1.0
Summary: Assertions for sensor-data pipelines. Every check aborts instead of warning.
Project-URL: Homepage, https://github.com/johnhagedorncs/sensorlint
Project-URL: Source, https://github.com/johnhagedorncs/sensorlint
Project-URL: Issues, https://github.com/johnhagedorncs/sensorlint/issues
Project-URL: Changelog, https://github.com/johnhagedorncs/sensorlint/blob/main/CHANGELOG.md
Author: John Hagedorn
Maintainer: John Hagedorn
License: MIT License
        
        Copyright (c) 2026 John Hagedorn
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Keywords: adc,assertions,data-pipeline,data-quality,sensor,signal-processing,telemetry,time-series,validation,vibration
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Software Development :: Quality Assurance
Classifier: Topic :: Software Development :: Testing
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: numpy>=1.22
Provides-Extra: dev
Requires-Dist: mypy>=1.8; extra == 'dev'
Requires-Dist: pytest-cov>=4.1; extra == 'dev'
Requires-Dist: pytest>=7.4; extra == 'dev'
Requires-Dist: ruff>=0.3; extra == 'dev'
Description-Content-Type: text/markdown

# sensorlint

Assertions for sensor-data pipelines. Every check aborts instead of warning.

```python
from sensorlint import safe_gunzip, assert_expected_length, assert_not_clipped

payload = safe_gunzip(raw_bytes)              # refuses to return partial data
samples = np.frombuffer(payload, dtype=np.int16)
assert_expected_length(samples, 20480)
assert_not_clipped(samples, adc_max=32767, leading_samples=2048)
```

---

## The thesis

Sensor pipelines do not usually fail by crashing. They fail by producing
believable numbers.

I built this because the bugs that cost me the most time all had the same
shape: nothing raised, nothing logged, the output was the right dtype and the
right shape and full of finite values in a plausible range, and it was wrong.
There was no traceback to work backwards from, and by the time anyone noticed,
the window for reproducing what had happened had closed.

Two examples, both of which have their own check in this library.

**A partial decode that one decoder calls an error and another calls a small
file.** `gzip.decompress` raises `EOFError` on a truncated stream. On the exact
same bytes, `zlib.decompressobj().decompress` returns whatever it managed to
inflate and does not raise — which is correct behaviour for that API, because it
has no way to know you are not about to feed it the rest of the stream. If any
layer of an ingest path falls back from the strict decoder to the permissive
one, truncated objects stop being errors and start being small files. A batch
job over a corrupted prefix then processes a fraction of what it thinks it
processed, or nothing at all, and exits zero. `examples/truncated_gzip.py`
demonstrates this end to end.

**A payload in raw converter counts whose scale factors live somewhere else.**
The counts are integers in a valid range. The spectrum is correct. The trend
over time is correct. Every relative comparison you make is correct. The values
are wrong by exactly one constant factor, and there is nothing to inspect,
because the output is plausible at any scale. A threshold of 0.4 g gets compared
against a value of 8,100 counts and fires, or does not fire, for reasons
unrelated to vibration.

The common property is that the wrong answer is indistinguishable from the right
one by looking at it. So the defence has to be an assertion made at the point
where the information still exists, not an inspection made later.

---

## Install

```bash
pip install git+https://github.com/<user>/sensorlint
```

Requires Python 3.10+ and numpy. Nothing else. (Not yet on PyPI; the wheel
builds and installs cleanly from a checkout.)

From a checkout:

```bash
make install     # creates .venv and installs with dev extras
make all         # ruff, mypy --strict, pytest
```

---

## 30-second quickstart

```python
import numpy as np
import sensorlint as sl

# 1. Decode without accepting partial data.
payload = sl.safe_gunzip(raw_bytes)
samples = np.frombuffer(payload, dtype=np.int16)

# 2. Assert what you were told to expect.
sl.assert_expected_length(samples, 20480)
sl.assert_not_clipped(samples, adc_max=32767, leading_samples=2048)
sl.assert_scale_declared(samples, scale=6.1e-5, unit="g")
sl.assert_not_stale(samples, max_frozen_run=64)
sl.assert_sample_rate(timestamps, claimed_fs=20_000, tolerance_pct=0.5)

# 3. At the end of every stage, the cheapest check in the library.
sl.assert_nonzero_kept(len(inputs), len(outputs), "featurize")
```

Every `assert_*` has a `check_*` twin that returns a structured result instead
of raising, so you can run the same checks over a fleet and get a report rather
than a traceback:

```python
results = [sl.check_not_clipped(load(f), adc_max=32767, target=f) for f in files]
print(sl.format_report(results))
```

```
sensorlint
==========

FAIL  1187/1204 checks passed (98.6%)

check        passed  failed  pass rate
-----------  ------  ------  ---------
not_clipped  1187    17      98.6%
TOTAL        1187    17      98.6%

failures by severity: error=17

worst offenders:
target          failures
--------------  --------
node-14/ch0     9
node-03/ch1     5
```

---

## The checks

| # | Check | Catches | Key option |
|---|-------|---------|------------|
| 1 | `assert_decoded_fully` / `safe_gunzip` | Truncated or partial decompression a permissive decoder swallowed | reads `decompressobj.eof`, not the data |
| 2 | `assert_expected_length` | Short reads, dropped samples, buffer truncation | `tolerance`, `tolerance_pct` |
| 3 | `assert_not_clipped` | Samples railed at the converter limit | `leading_samples` — the variant that matters |
| 4 | `assert_scale_declared` | Raw ADC counts with no scale factor and unit | `looks_like_raw_counts` heuristic |
| 5 | `assert_not_stale` | Frozen sensors, and linearly interpolated stretches posing as measurements | `max_interpolated_run` |
| 6 | `assert_sample_rate` | A claimed rate that the timestamps contradict; gaps; jitter | `gap_factor`, `allow_gaps` |
| 7 | `detect_channel_transposition` | Swapped channel pairs, plus a **false-positive rate with no ground truth** | `transposition_run_lengths` |
| 8 | `assert_coverage` | Data *around* a date range rather than *on* it | `bin_seconds`, `min_density` |
| 9 | `assert_nonzero_kept` | A stage that processed zero records and reported success | `@nonzero_kept`, `pipeline_stage` |

---

## Design principles

**1. Abort, never warn.** There is no warning mode and there will not be one. A
warning in a batch job is a log line nobody reads, and the entire premise of
this library is that these failures are already invisible. Every `assert_*`
raises. `SensorLintError` subclasses `AssertionError`, but unlike the `assert`
statement it is not removed by `python -O`, which matters because production
batch jobs are where these checks earn their keep.

**2. Every check exists twice.** The `check_*` form returns a `CheckResult`
dataclass — check name, pass/fail, severity, message, and a `details` dict of
the numbers behind the verdict. The `assert_*` form is a thin wrapper that calls
it and raises. This is not decoration: a single file failing a check is an
exception, and ten thousand files failing a check is a report, and you cannot
build a report out of tracebacks. `sensorlint.report` aggregates results by
check, by severity and by target.

**3. Zero heavy dependencies.** numpy and the standard library. Nothing else,
ever. A validation library that is annoying to install is a validation library
that gets skipped at exactly the ingest boundary where it was supposed to run.

**4. Every check has a test that constructs the real failure.** Not a mock. The
gzip tests truncate actual compressed bytes and confirm that `zlib` tolerates
what `gzip` rejects. The clipping tests actually saturate arrays. The
transposition tests actually swap two real signal arrays with different energy.
The staleness tests actually overwrite a stretch of a noisy signal with
`np.linspace` between its endpoints. If the failure cannot be constructed, I do
not trust the check to detect it.

---

## The checks in detail

### 1. Decoded fully

The tell is not the data, it is the decompressor state. A complete stream leaves
`decompressobj.eof` set to `True`; a truncated one leaves it `False` with nothing
in `unused_data`. That single boolean is the whole check. `probe_compressed`
re-decodes strictly and reports `eof`, member count, unconsumed tail, and the
gzip CRC32/ISIZE trailer; `check_decoded_fully` compares that against what your
decoder handed you.

`safe_gunzip` is the function to reach for instead of `gzip.decompress` or a
hand-rolled `decompressobj` loop. It verifies end-of-stream before returning
anything, so a truncated object raises rather than becoming a short buffer that
looks like a small file. Multi-member gzip and trailing NUL padding are handled;
neither is corruption.

### 2. Expected length

The most boring check here and one of the most common failures. A short read
gives you an array of entirely real measurements — of the first 200 ms of a
one-second record. Your RMS is a real RMS. Your spectrum has a fifth of the
frequency resolution you documented. Nothing downstream can tell. Write the
expected length down at the point where you still know it.

### 3. Not clipped

A saturated converter returns its ceiling, repeatedly, while the input stays out
of range. The result has a *lower* peak-to-peak than the truth and a *higher*
apparent harmonic content, so it flatters every summary statistic you compute.

The reason this is more than a one-liner is the leading window. Railing very
often happens only at the start of a recording, while a coupling capacitor
charges or a gain stage settles, and then decays into perfectly normal values.
Over a 60-second record, 200 ms of solid rail is 0.3% of the samples — under any
whole-array threshold you would actually set. So the first block of every file is
garbage, every file, forever, and the per-file statistics look fine. Pass
`leading_samples=N` and the window is tested separately.

When `adc_max` is not supplied, rails are auto-detected from *repeated exact*
extreme values. Noisy data reaches its maximum once; a saturated converter
returns bit-identical extremes many times over. The signature is the repeat
count, not the magnitude.

### 4. Scale declared

Refuses to proceed on data that looks like raw counts with no declared scale
factor and physical unit. `looks_like_raw_counts` scores several signals:
integer dtype, float values that are all whole numbers (the usual way counts get
past a type check), an extreme sitting just under a common converter full scale,
and a quantization step of exactly 1.

It is deliberately conservative in one direction. It would rather make you
declare a unit for data that was already in physical units than let undeclared
counts through, because there is no way to detect the second mistake afterwards.

### 5. Not stale

Frozen data has zero variance over the stuck stretch, which reads to most
anomaly detectors as "very stable". Interpolated data hides better, because it
still moves.

The discriminant for interpolation is the second difference. A linearly
interpolated stretch has a constant first difference, so its second difference is
zero to floating-point precision, across many consecutive samples. Real sensor
data essentially never does this — even a very clean, very slowly varying signal
has noise in the last bits, and that is enough. A long run of `d2 == 0` is not a
property of quiet data, it is a signature of synthesis. Zero slope is reported as
frozen; constant non-zero slope is reported as interpolated.

### 6. Sample rate

The sampling rate is nearly always metadata; the timestamps are data. A spectrum
computed with a claimed 20 kHz on data actually sampled at 19,531 Hz puts every
peak 2.4% low, so a fault frequency lands next to where you were looking rather
than on it. The plot is beautiful.

The check compares the claim against the median observed interval (median, so
one dropped block does not move the estimate), reports jitter, and refuses to
call a record contiguous when it has holes. Concatenating across a hole produces
a discontinuity that reads as broadband energy — indistinguishable from an
impact event, which is often exactly what the pipeline is looking for.

### 7. Channel transposition — the one I would want to talk about

Two channels get swapped: a connector goes back the wrong way, a channel map is
edited, a wiring change is made and the metadata is not. Both channels still
contain real signal. Every per-channel statistic afterwards is a correct
measurement of the wrong thing, so a rising trend appears on the wrong location
and the one that is actually degrading looks stable.

The detector uses the inter-channel energy ratio, scored as two hypotheses
rather than one threshold. If A normally runs 8 dB above B, then a file where A
runs 8 dB *below* B is not "a quiet day on A". The verdict compares the distance
from the observed ratio to `+expected` against the distance to `-expected`, and
requires a margin. Ratios in the ambiguous middle return `transposed=False` with
low confidence, which is the honest answer, and channels whose expected
separation is too small to discriminate raise rather than emit coin flips.

That detector has a false-positive rate and no ground truth to measure it
against. Nobody is going to physically inspect a thousand installations to tell
you how often you were wrong.

But transposition has a property noise does not: **it is temporally contiguous.**
A cable that got swapped stays swapped until someone unplugs it. So across a
chronologically ordered sequence of files from one installation, real
transposition appears as a *run* — dozens or hundreds of consecutive positives —
while a detector firing on noise produces isolated single-file positives
scattered through long stretches of negatives.

That gives you an error rate for free:

```python
verdicts = [detect_channel_transposition(a, b, expected_ratio_db=8.0) for a, b in files]
report = transposition_run_lengths(verdicts, min_run_length=5)

report.contiguous_runs                    # ((150, 40),) — the actual wiring change
report.n_isolated                         # 3 — lone flags, believed spurious
report.estimated_false_positive_rate      # 3 / 360 = 0.0083, with no labels
report.n_isolated_expected_if_random      # what pure noise would have produced
report.clustered                          # True: these flags carry real structure
```

`estimated_false_positive_rate` is the count of positives sitting in short runs
divided by the number of files believed to be true negatives.
`n_isolated_expected_if_random` is the same count under an i.i.d. Bernoulli model
at the observed flag rate — if your detector's isolated count matches it, the
positives have no temporal structure and are consistent with pure chance.

The same structure tells you when to believe a positive. One flagged file in a
year of clean ones is your noise floor. Forty consecutive flagged files starting
the week after a site visit is a wiring change.

Use `estimate_expected_ratio_db(reference_pairs)` to learn the baseline from
known-good files. It takes the median, so a single already-swapped file in the
reference set cannot drag the baseline toward zero and quietly disarm the check.

### 8. Coverage

You ask for a date range. Something returns files whose timestamps fall inside
it. The earliest is at or before the start, the latest is at or after the end,
both are true, and you conclude the range is covered.

It is not. A file on the first day and a file on the last day satisfy every
endpoint test ever written while leaving a month-long hole in between. The
statistics you compute over "the last 90 days" are computed over two days, the
trend line has two points, and the baseline is a baseline of nothing.

This check bins the range and requires the bins to be populated. Endpoint
coverage is necessary and nowhere near sufficient. It reports the longest run of
empty bins, so the failure message names the outage rather than just its
existence.

### 9. Nonzero kept

**This is the cheapest assertion in the library and it catches the most
expensive class of bug.**

A stage iterates over its inputs, filters, transforms, writes its outputs. The
list was empty, or every record was filtered out by a predicate that stopped
matching after an upstream schema change. The loop body never runs. There is
nothing to raise. The stage writes an empty output, logs "stage complete", and
exits zero. Everything downstream then works perfectly on nothing. The job is
green. The dashboard is green.

It is one comparison against zero. It needs no thresholds, no domain knowledge,
and no numerical thinking, and it costs nothing at runtime. And a silent empty
result can survive for weeks, because it does not look like a failure — it looks
like a quiet period. By the time somebody notices the gap, the logs have rotated
and the window for working out what went wrong has closed.

Put it at the end of every stage. Not the important ones. Every one. Three forms
are provided so there is never a reason to skip it:

```python
assert_nonzero_kept(len(inputs), len(outputs), "featurize")

@nonzero_kept("decode_batch")
def decode_batch(objects): ...

with pipeline_stage("decode", n_input=len(objects)) as stage:
    for obj in objects:
        stage.keep()
```

The context manager does not mask an exception raised inside the block: a stage
that crashed already failed loudly, and replacing its traceback with an
empty-output complaint would hide the real cause.

`examples/silent_empty_stage.py` runs a three-stage pipeline before and after an
upstream field rename, with and without the assertion.

---

## Command line

```bash
python -m sensorlint check recording.npy \
    --expected-length 20480 \
    --adc-max 32767 --leading-samples 2048 \
    --scale 6.1e-5 --unit g

python -m sensorlint check series.csv \
    --column 1 --timestamp-column 0 --fs 20000 --json
```

Exits `0` when every applicable check passed, `1` when any failed, `2` on a
usage or loading error. Checks with no corresponding flag are skipped rather
than run against a guessed threshold, and skipped checks are printed explicitly
so they cannot hide inside a green run.

---

## Errors

One exception class per check family, so a caller who genuinely wants to
tolerate one class of problem can do so narrowly instead of writing
`except Exception` and thereby also swallowing the problem that matters.

```
SensorLintError(AssertionError)
├── DecodeError          ├── StalenessError
├── LengthError          ├── SampleRateError
├── ClippingError        ├── TranspositionError
├── ScaleError           ├── CoverageError
                         └── EmptyStageError
```

Every exception carries the `CheckResult` that produced it, so the structured
detail survives the raise:

```python
try:
    assert_not_clipped(samples, adc_max=32767, leading_samples=2048)
except ClippingError as exc:
    exc.details["leading_railed_fraction"]   # 0.37
    exc.result.target                        # "node-14/ch0"
```

---

## Notes

**On the Python version.** The package declares `requires-python = ">=3.10"`,
but every module carries `from __future__ import annotations` and avoids
3.10-only syntax at runtime, so the source also imports cleanly on 3.9. The ruff
config disables the pyupgrade rules that would break that. This is deliberate:
ingest code is often the last thing on an old interpreter, and a validation
library that cannot be installed there is a validation library that does not run.

**Where this belongs.** At the ingest boundary: the first thing that touches a
payload after it is read, before anything computes on it. Wiring it into the
ingest stage of my [`bearing-watch`](https://github.com/johnhagedorncs/bearing-watch)
project is next.

**Everything in this repository is synthetic.** All example data, thresholds and
signals are generated in-repo. No proprietary code, data or results from any
employer appear anywhere in it.

---

## Development

```bash
make install     # .venv + dev extras
make test        # pytest
make test-cov    # pytest with coverage
make lint        # ruff
make typecheck   # mypy --strict
make examples    # run both example scripts
make all         # lint + typecheck + test
```

464 tests, 97% branch coverage, mypy strict clean. CI runs ruff, mypy strict, and
pytest on 3.10 / 3.11 / 3.12.

---

## License

MIT. See [LICENSE](LICENSE).
