Metadata-Version: 2.4
Name: shardorm
Version: 0.0.5
Summary: A lightweight Python micro-ORM with automatic PostgreSQL sharding, replication, failover, and schema synchronization.
License: AGPL-3.0
Classifier: Programming Language :: Python :: 3
Classifier: License :: OSI Approved :: GNU Affero General Public License v3
Classifier: Operating System :: OS Independent
Requires-Python: >=3.11
Description-Content-Type: text/markdown
Requires-Dist: psycopg[binary]>=3.0.0
Requires-Dist: psycopg_pool>=3.0.0
Provides-Extra: dashboard
Requires-Dist: fastapi>=0.100.0; extra == "dashboard"

# ShardORM

> A lightweight Python micro-ORM with automatic PostgreSQL sharding, replication, failover, and schema synchronization.
>
> **Zero cluster management. Zero coordination layers. Pure Python over standard Postgres.**

ShardORM is built directly on top of **psycopg 3**, **psycopg_pool**, and the Python standard library. It automatically distributes data across PostgreSQL shards while keeping your schema synchronized on every node.

---

## Features

🗄️ Automatic PostgreSQL sharding
Deterministically distributes records across multiple PostgreSQL database shards without requiring manual routing logic in your application code.

🔄 Consistent hashing with virtual nodes
Uses a consistent hash ring with virtual nodes (CRC32), ensuring that adding or removing shards only requires migrating a minimal fraction of keys.

📦 Configurable replication factor
Allows flexible definition of how many shards a record should be replicated to for fault tolerance (e.g., replicating across 2 or 3 shards).

⚡ Automatic read failover
Queries target shards sequentially when the shard key is present, falling back automatically until the first responsive shard succeeds.

🌍 Full table replication (full_sync)
Enables complete synchronization of smaller tables (such as global settings or lookup data) across all configured shards.

🔀 Zero-downtime online rebalancing
Supports scaling and decommissioning shards (draining_shards) with live background data migration via dual-writes and routing layers (migrate-online).

📜 SQL-based migrations
Keeps database schemas strictly synchronized across all shards—either via ad-hoc commands or versioned .sql migration files (acting like a mini-Alembic).

🛡️ Zero-downtime DDL (add-column-safe & enforce-not-null)
Executes schema modifications safely under production load using short lock_timeout retry loops and deferred constraints (CHECK ... NOT VALID) to prevent table locks.

🔒 Optional distributed transactions (PostgreSQL 2PC)
Enables true distributed "all-or-nothing" writes across multiple shards using PostgreSQL's native Two-Phase-Commit (PREPARE TRANSACTION). Backed by an automated recovery daemon that cleans up orphaned locks after crashes.

🔄 Distributed coordinator state
Maintains 2PC decisions, saga trails, and migration states inside dedicated tables on the PostgreSQL shards themselves—enabling a completely stateless multi-instance application setup.

🎭 Saga Pattern (compensating transactions)
Orchestrates multi-shard business logic with automatic rolling back of completed steps via custom compensating actions if a step fails.

🏊 Built-in connection pooling
Leverages performant and robust per-shard connection management via psycopg_pool, including automatic timeouts and connection cleanup.

🛡️ Built-in SQL Injection Protection
Validates all table and column names strictly against a whitelist regular expression and safely masks them via psycopg.sql.Identifier.

🛑 Built-in Circuit Breaker
Detects failing shards in real time, immediately short-circuiting connection attempts during a cooldown period to prevent thread pool exhaustion.

⚡ Parallel scatter-gather queries
Executes global and multi-shard read operations concurrently using an optimized thread pool (ThreadPoolExecutor) to keep response latencies minimal.

📊 Built-in FastAPI Dashboard
Provides an optional, drop-in APIRouter featuring a real-time dark-mode web UI and JSON API to inspect shard health, connection pools, circuit breakers, cluster consistency, and trigger migrations or recovery tasks.

---

## Installation

```bash
pip install shardorm fastapi  # FastAPI is optional, only required for the dashboard
```

Requirements:

- Python 3.11+
- PostgreSQL 13+
- psycopg 3
- psycopg_pool

---

## Configuration

ShardORM loads its configuration from:

```
./shardorm.config.json
```

or

```
$SHARDORM_CONFIG
```

Generate a template:

```bash
shardorm init-config
```

Example:

```json
{
  "replication_factor": 2,
  "min_write_quorum": 1,
  "shards": [
    "postgresql://user:password@localhost:5432/appdb1",
    "postgresql://user:password@localhost:5432/appdb2",
    "postgresql://user:password@localhost:5432/appdb3"
  ],
  "table_policies": {
    "users": {
      "mode": "sharded",
      "shard_key": "id",
      "replication_factor": 2,
      "write_mode": "quorum"
    },
    "orders": {
      "mode": "sharded",
      "shard_key": "id",
      "replication_factor": 2,
      "write_mode": "2pc"
    },
    "categories": {
      "mode": "full_sync"
    }
  }
}
```

---

## Table Policies

### Sharded

Rows are distributed across the cluster using a consistent hash ring.

```json
{
  "users": {
    "policy": "sharded",
    "shard_key": "id"
  }
}
```

### Full Sync

Rows are replicated to every configured shard.

```json
{
  "countries": {
    "policy": "full_sync"
  }
}
```

---

## CLI

```bash
# Configuration & Status
shardorm init-config
shardorm status

# Schema & Migrations
shardorm make-migration add_index_users_email
shardorm migrate
shardorm add-table users id:UUID:PK email:TEXT data:JSONB
shardorm add-column users status:TEXT

# Zero-Downtime DDL (Under Production Load)
shardorm add-column-safe users role:TEXT
shardorm enforce-not-null users role

# Shard Rebalancing (Offline vs Online)
shardorm rescale users              # Dry-Run preview
shardorm rescale users --apply      # Blocking block-copy
shardorm migrate-online users --start
shardorm migrate-online users --worker-loop
shardorm migrate-online users --cutover

# Distributed Transactions Recovery
shardorm recover-2pc
shardorm 2pc-daemon --interval 30
```

---

## FastAPI Example

```python
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from shardorm import ShardORM

app = FastAPI(title="ShardORM Example")

db = ShardORM.from_config()


class UserCreate(BaseModel):
    id: str
    name: str
    email: str


@app.on_event("shutdown")
def shutdown():
    db.close()


@app.get("/users")
def get_all_users():
    try:
        users = db.table("users").select_all_shards()
        return {
            "status": "success",
            "count": len(users),
            "users": users
        }
    except Exception as e:
        raise HTTPException(500, str(e))


@app.post("/users")
def create_user(user: UserCreate):
    try:
        result = (
            db.table("users")
            .insert(
                id=user.id,
                data=user.model_dump()
            )
        )

        return {
            "status": "success",
            "write_result": result
        }

    except Exception as e:
        raise HTTPException(500, str(e))


@app.get("/users/{user_id}")
def get_user(user_id: str):
    users = (
        db.table("users")
        .where(id=user_id)
        .select()
    )

    if not users:
        raise HTTPException(404, "User not found")

    return users[0]


@app.get("/cluster/status")
def cluster_status():
    return {
        "shards": db.status()
    }
```

## FastAPI Dashboard Integration

```python
from fastapi import FastAPI, Depends
from shardorm import ShardORM, create_dashboard_router

app = FastAPI(title="ShardORM App")
db = ShardORM.from_config()

# Mount dashboard under /shardorm
app.include_router(
    create_dashboard_router(db),  # optional: dependencies=[Depends(require_admin)]
    prefix="/shardorm"
)

@app.on_event("shutdown")
def shutdown():
    db.close()
```

## Saga Pattern Example (Multi-Shard Workflows)

```python
from shardorm import ShardORM, SagaStep

def debit_account(db):
    db.table("accounts").where(id="acc_1").update(balance=500)

def credit_account(db):
    db.table("accounts").where(id="acc_2").update(balance=1500)

def rollback_debit(db):
    db.table("accounts").where(id="acc_1").update(balance=600)

db = ShardORM.from_config()

try:
    db.saga([
        SagaStep("debit_source", action=debit_account, compensation=rollback_debit),
        SagaStep("credit_target", action=credit_account),
    ]).run()
except Exception as e:
    print(f"Saga failed and automatically compensated: {e}")
```

---

## How It Works

```
Shard Key
     │
     ▼
Hash Function
     │
     ▼
Consistent Hash Ring
     │
     ▼
Virtual Nodes
     │
     ▼
Replication Factor
     │
     ▼
Destination Shards
```

Only the required shards are contacted for reads and writes. When new shards are added, only the affected rows are moved during rebalancing.

---

## Two-Phase Commit (Optional)

Enable atomic distributed transactions for a table:

```json
{
  "orders": {
    "policy": "sharded",
    "write_mode": "2pc"
  }
}
```

Internally this uses PostgreSQL's native:

```sql
PREPARE TRANSACTION
COMMIT PREPARED
ROLLBACK PREPARED
```

> PostgreSQL requires `max_prepared_transactions > 0`.

---

## Why ShardORM?

| Feature | ShardORM |
|----------|-----------|
| ORM | Lightweight |
| Pure SQL | ✅ |
| psycopg 3 | ✅ |
| Connection Pooling | ✅ |
| Sharding | ✅ |
| Replication | ✅ |
| Read Failover | ✅ |
| Online Rebalancing | ✅ |
| SQL Migrations | ✅ |
| PostgreSQL 2PC | ✅ |

---

## License

Licensed under the **GNU Affero General Public License v3.0 (AGPL-3.0)**.
