Metadata-Version: 2.5
Name: reble
Version: 0.6.1
Summary: Git-style branching for Iceberg data warehouses: scoped branches, row-level diffs, fast-forward promotion.
Author: Reble contributors
License: Apache-2.0
License-File: LICENSE
Keywords: branching,data-engineering,dbt,iceberg,lakehouse
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Database
Requires-Python: >=3.10
Requires-Dist: duckdb>=1.0
Requires-Dist: httpx>=0.27
Requires-Dist: pydantic>=2.7
Requires-Dist: pyiceberg[pyarrow,sql-sqlite]>=0.8
Requires-Dist: pyyaml>=6.0
Requires-Dist: rich>=13.0
Requires-Dist: sqlglot>=25.0
Requires-Dist: typer>=0.12
Provides-Extra: aws
Requires-Dist: pyiceberg[glue]>=0.8; extra == 'aws'
Provides-Extra: dev
Requires-Dist: mypy>=1.10; extra == 'dev'
Requires-Dist: numpy>=1.26; extra == 'dev'
Requires-Dist: pre-commit>=3.7; extra == 'dev'
Requires-Dist: pytest-cov>=5.0; extra == 'dev'
Requires-Dist: pytest>=8.0; extra == 'dev'
Requires-Dist: ruff>=0.5; extra == 'dev'
Provides-Extra: mcp
Requires-Dist: mcp>=2.1; extra == 'mcp'
Provides-Extra: postgres
Requires-Dist: psycopg2-binary>=2.9; extra == 'postgres'
Provides-Extra: spark
Requires-Dist: pyspark<3.6,>=3.5; extra == 'spark'
Description-Content-Type: text/markdown

# Reble

[![CI](https://github.com/satya1395/reble/actions/workflows/ci.yml/badge.svg)](https://github.com/satya1395/reble/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/reble)](https://pypi.org/project/reble/)
[![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)

*Reble* — pronounced **re-bl** (the final *e* is silent).

**Reble is an open SQL engine for your Iceberg lakehouse.** Your models are
plain SQL files; Reble derives their dependencies, builds the tables into
your catalog and bucket, and refreshes exactly what moved — triggered by
cron, CI, Airflow, or an agent.

Every data team ends up building the same expensive hack around that job:
a copy of the warehouse for testing changes. You refresh it, you queue for
it, you hope it still matches prod — and you pay for it twice. The Apache Iceberg
table format already has the primitive that makes it unnecessary: a branch
ref is
metadata-only, so a "copy" of a 5M-row table costs **10 ms and zero
bytes** on any compliant catalog. What was missing is the *workflow* —
deciding what a change touches, making its inputs reproducible, showing
what it will do to production rows before anyone accepts it.

Reble is that workflow, and it isn't bolted on: scope, pin, run, diff,
promote are what the engine does. When you change a model, it branches,
re-runs only the blast radius — your edited models plus their downstream
closure, derived from the SQL — shows you the exact rows that will change,
and fast-forwards production when you accept. There is no merge step, on
purpose: promote or discard, never three-way-merge data, because data
merges are where correctness goes to die.

![The Reble loop: build, edit on a branch, diff the rows, promote](docs/assets/demo.gif)

## How it works

```mermaid
flowchart TB
    WHO["who triggers — cron · CI · Airflow · AI agents (MCP)"]
    MODELS["your models — models/*.sql, plain SQL + a 3-line header"]
    REBLE["Reble — SQLGlot lineage · scope · pin · run · diff · promote"]
    ENGINE["compute — DuckDB (default) · Spark (same interface)"]
    CAT["your Iceberg catalog — Glue · Polaris · Nessie · Hive · REST · sql"]
    STORE[("your storage — S3 · GCS · local disk")]
    WHO -->|"invokes one verb"| REBLE
    MODELS --> REBLE
    REBLE --> ENGINE
    ENGINE -->|"branch refs · tag pins · snapshots"| CAT
    CAT --> STORE
```

Reble owns the *transformation* layer — models, lineage, execution,
branching — the shape dbt-core has, without the templating or YAML. It
does **not** own *scheduling*: cron or Airflow decides when; Reble is the
step they run. And it's built on **native Iceberg branch refs** — a
per-table Iceberg spec feature supported by any catalog (Glue, Polaris,
Nessie, Hive, or any REST-compliant catalog). It is *not* a catalog and
requires no new infrastructure. A branch ref is metadata-only: zero bytes
are copied.

## Quick start

```
pip install reble
reble init                # writes reble.yml; probes your catalog
git switch -c fix-orders  # or: --change-set agent-42 — git is one adapter
# ...edit two models...
reble run                 # → data branch: edited models + downstream closure
                          #   written; upstream inputs pinned via Iceberg tags
reble diff                # schema + row-level diff vs. branch base
reble status              # un-run edits, drifted pins, branch age/expiry
reble promote             # fast-forward if base is current; forced re-run with
                          #   fresh diff if main moved. No merge. Ever.
```

## What Reble is — and isn't

- **Is:** a transformation engine (models + lineage + execution) that works
  with the Iceberg catalog you already run (Glue, Polaris, Nessie, Hive,
  any REST catalog). No server, no new infrastructure.
- **Isn't:** a scheduler (cron/Airflow's job — Reble is the step they run),
  a catalog, or a merge tool. There is no
  three-way data merge, ever — a change is either fast-forwarded or re-run.
- See [how Reble compares](https://satya1395.github.io/reble/comparisons/)
  to lakeFS, Nessie, and warehouse clones.

## The loop in detail

```mermaid
flowchart LR
    M[("main<br/>(Iceberg tables)")]
    E["edited SQL"] -->|"scope: AST-changed ∪<br/>downstream closure"| RUN["reble run"]
    M -->|"upstream inputs pinned<br/>via Iceberg tags"| RUN
    RUN -->|"zero-copy branch refs"| B[("data branch")]
    B --> D["reble diff<br/>rows + schema"]
    D --> P{"reble promote"}
    P -->|"pinned bases still<br/>equal main"| FF["fast-forward main"]
    P -->|"drift"| RR["scoped re-run +<br/>fresh promote-time diff"]
    RR --> FF
```

- **Scoped branching** — scope = edited models ∪ downstream closure, capped
  by `--depth`; `reble run --refresh` scopes by *data* movement instead
  (nightly refreshes rebuild exactly what ingested).
- **Pinned inputs** — upstream tables pinned with Iceberg **tags**
  (`reble_pin__*`) at run time; tags block `expire_snapshots`, so branch
  reads stay correct even while main moves.
- **Row-level diffs** — computed on your compute via DuckDB, streaming
  through `iceberg_scan` (out-of-core; spills under a configurable
  `engines.duckdb.memory_limit`).
- **Promote semantics** — fast-forward only when every pinned base still
  equals current main; otherwise a scoped re-run and a fresh, promote-time
  diff. The PR diff is advisory; the promote diff is authoritative.
- **Atomic per-table commits** — every model write and every promotion
  step commits a table atomically: no partial states, no half-applied
  changes, interrupted work resumes instead of repeating. Cross-table
  consistency is *verified at promote time* (every pin still equals
  main) — not reconciled afterward by a merge you have to trust.

## Performance

Measured, reproducible, no clusters — full numbers and reproduction
commands on the [performance page](https://satya1395.github.io/reble/performance/):

- **Branch a 5M-row table: < 10 ms.** Branches are metadata-only.
- **Full lifecycle on AWS (Glue + S3, 1M rows, from a laptop):** scoped
  run ~13 s, keyed diff ~4 s, drift check ~2 s — with streaming reads
  verified engaged (`iceberg_scan`, zero fallbacks).
- Reads spill under a configurable `memory_limit` — working set bounded by
  config, not by RAM.

## Models are plain SQL

**"Model" is just our word for one SQL file that creates one table.** If
your team keeps a folder of SQLs and schedules them some way — a DAG, cron,
an internal webapp — you already have models; point Reble at the folder.
No orchestrator required, no dbt required: `models/**/*.sql`, one file is
one model, the file stem is the table name, and a minimal header comment
block carries the semantics:

```sql
-- model: mart_orders      (optional; defaults to file name)
-- kind: table | view
-- key: order_id           (diff key)
select ... from stg_orders join raw_customers using (customer_id)
```

Every run fully rebuilds its scope — replace, never append — and
`reble run --force` rebuilds even unchanged SQL. (There is deliberately no
`incremental` kind yet: it arrives when watermark / insert-overwrite
execution is real, not as a word that recomputes everything.)

Lineage is parsed with SQLGlot: a table reference that matches another model
is an edge; anything else is an upstream input, pinned with an Iceberg tag at
run time. Cosmetic edits (whitespace, comments, casing) hash identically on
the canonical AST and never trigger a run. Every branch snapshot carries
provenance (`reble.model`, `reble.ast_hash`, `reble.run_id`) in its summary —
"which code produced this table state" is answered from the catalog itself.

## Set it up once, it runs itself

No manual runs. Put one command in cron, CI, or Airflow — the same
command whether you have two models or two hundred:

```bash
# nightly: rebuild exactly what changed
0 3 * * * cd /srv/warehouse && reble run --refresh && reble gc
```

Safe to trigger repeatedly: a night with no new data does nothing. A
crashed run resumes from where it stopped. A retried promote never
double-applies.

For teams, state lives in Postgres so multiple workers share it safely
(one line in `reble.yml`, validated at startup — no shared filesystem
needed):

```yaml
state:
  store: postgres
  uri: ${REBLE_STATE_URI}   # postgresql://user:pass@host:5432/reble
```

Install with `pip install 'reble[postgres]'`.

The verbs are idempotent and the exit codes are a contract, so any job can
drive Reble: scheduled, or triggered by whatever lands your data. A double
trigger is harmless — a quiet night computes an empty scope from one
catalog listing.

```yaml
# .github/workflows/refresh.yml — rebuild exactly what moved, nightly and
# on demand (your ingestion job can dispatch it when new data lands)
on:
  schedule: [{cron: "0 3 * * *"}]
  workflow_dispatch:
jobs:
  refresh:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: pip install reble
      - run: reble run --refresh     # + reble gc to drop expired branches
        env: { REBLE_CHANGE_SET: local }   # catalog/warehouse creds as secrets
```

PRs get the same treatment: `reble status` exits 3 on drift and
`reble diff` prints the row-level consequences — cheap checks to wire into
any CI.

## Agents (MCP)

Any MCP host can drive the same verbs — the agent has no special powers:

```json
{
  "mcpServers": {
    "reble": {
      "command": "reble-mcp",
      "env": { "REBLE_PROJECT_DIR": "/path/to/project" }
    }
  }
}
```

Install with `pip install 'reble[mcp]'`. `reble_run` generates and returns a
change-set id; errors carry the spec exit codes as structured `error.code`
(3 = drift, 4 = promote-blocked). Tool docstrings are the agent-facing spec.

Agents and CI are first-class everywhere, not just over MCP: every command
speaks a stable [`--json` envelope](SPEC.md) with documented exit codes,
`run`/`diff` stream versioned [`--events`](SPEC.md#event-streams) (NDJSON),
and work is keyed by change-set (`--change-set <id>` or `REBLE_CHANGE_SET`)
so it never depends on git — `--branch` resumes an existing data branch
under a new change-set.

## Documentation

- [**Docs site**](https://satya1395.github.io/reble/) — getting started,
  concepts, comparisons, and the command reference.
- [`SPEC.md`](SPEC.md) — normative CLI specification (v0.2): invariants,
  on-disk layout, `reble.yml` schema, command reference, JSON envelope,
  event streams, provenance, exit codes.
- [`DECISIONS.md`](DECISIONS.md) — recorded behavior decisions.

## Requirements

- Python 3.10–3.13 (3.14 not yet tested)
- An S3 bucket + an Iceberg catalog (Glue, Polaris, Nessie, Hive, or any
  REST-compliant one) — on AWS, `pip install 'reble[aws]'` and follow the
  [AWS walkthrough](https://satya1395.github.io/reble/aws/) (covers bucket
  creation, credentials, and every step from zero)
- SQL models under `models/` (path configurable via `lineage.models_path`)

## Roadmap to 1.0

| Release | Theme | Highlights | Status |
| --- | --- | --- | --- |
| 0.4 | Runs on AWS | Glue + S3 verified end-to-end, credential auto-config, self-cleaning AWS smoke | ✅ shipped |
| 0.5 | Bigger warehouses | Spark runner (local first, then serverless); GCS + ADLS verification; partitioned tables; incremental execution (watermark / insert-overwrite) | Spark runner ✅ (0.6.0); rest ongoing |
| 0.6 | Backfills & teams | Date-range / partition-scoped backfills (branch + insert-overwrite); documented CI recipes (PR checks, promote gates); multi-writer etiquette; `reble doctor` | planned |
| 0.7 | Interop | REST catalogs verified (Polaris, Nessie); Trino read adapter on demand; Iceberg views | planned |
| 0.8 | Operations | Metrics/log hooks; `estimate` v2; Windows support | planned |
| **1.0** | **GA** | See criteria below | — |

**GA criteria** — 1.0 ships when, not before: the JSON envelope, event
streams, and exit codes have held stable through a full minor release;
the lifecycle is green in CI on Glue + one REST catalog + sql; both
engines (DuckDB, Spark) are real; at least three non-trivial warehouse
deployments have run promote in production; and a security pass is done.

## License

Apache-2.0.
