Metadata-Version: 2.4
Name: hbase-driver
Version: 1.1.1
Summary: Pure-Python native client for Apache HBase, without Thrift.
Project-URL: Homepage, https://github.com/innovationb1ue/hbase-driver
Project-URL: Documentation, https://innovationb1ue.github.io/hbase-driver/
Project-URL: Source, https://github.com/innovationb1ue/hbase-driver
Project-URL: Changelog, https://github.com/innovationb1ue/hbase-driver/blob/main/docs/CHANGELOG.md
Project-URL: Issues, https://github.com/innovationb1ue/hbase-driver/issues
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: kazoo
Requires-Dist: protobuf
Dynamic: license-file

# hbase-driver

**English | [简体中文](https://github.com/innovationb1ue/hbase-driver/blob/main/README_CN.md) | [Documentation](https://innovationb1ue.github.io/hbase-driver/)**

[![Tests](https://github.com/innovationb1ue/hbase-driver/actions/workflows/ci.yml/badge.svg)](https://github.com/innovationb1ue/hbase-driver/actions)
[![License](https://img.shields.io/badge/license-Apache%202-blue.svg)](https://github.com/innovationb1ue/hbase-driver/blob/main/LICENSE)

`hbase-driver` is a pure-Python native client for Apache HBase 2.x. It does not require Thrift: the driver communicates directly with HBase RegionServers and the Master over RPC and uses ZooKeeper for service discovery.

## Features

- Native HBase RPC access without a Thrift Server
- Core data operations: `Put`, `Get`, `Delete`, and `Scan`
- Batch reads and writes, atomic increments, conditional mutations, appends, and row mutations
- Server-side filters, connection pooling, and Region location caching
- Table, column family, namespace, snapshot, Region, and cluster administration
- Context-manager support for `Client`, `Table`, `Admin`, `ResultScanner`, and `BufferedMutator`
- An API style familiar to users of the HBase Java Client

## Requirements

- Python 3.10 or newer is recommended; the development container uses Python 3.11
- A reachable HBase 2.x cluster
- Network access to ZooKeeper and to the HBase Master and RegionServer addresses returned by ZooKeeper
- The provided Docker integration environment uses HBase 2.6.1

> The client discovers HBase services through ZooKeeper and currently uses the default `/hbase` znode. No Thrift service is required.

## Installation

Install directly from PyPI:

```bash
python -m pip install hbase-driver
```

Creating a virtual environment is optional. The installation command above works in an existing Python environment. The `kazoo` and `protobuf` runtime dependencies are installed automatically.

### Optional: use a virtual environment

Use a virtual environment if you want to isolate this package and its dependencies from other Python projects:

```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install hbase-driver
```

On Windows PowerShell, activate it with:

```powershell
.venv\Scripts\Activate.ps1
```

### Install from source

```bash
git clone https://github.com/innovationb1ue/hbase-driver.git
cd hbase-driver
python -m pip install .
```

For development, install the project in editable mode together with its test dependencies:

```bash
python -m pip install -e .
python -m pip install -r requirements.txt
```

### Verify the installation

```bash
python -c "import hbasedriver; print('hbase-driver installed successfully')"
```

This command only verifies that the package can be imported; it does not connect to an HBase cluster.

## Quick start

The following example assumes that the `default:mytable` table and its `cf` column family already exist:

Use a disposable table: this example overwrites a cell and deletes the entire `row1` row.

```python
from hbasedriver.client.client import Client
from hbasedriver.hbase_constants import HConstants
from hbasedriver.operations.delete import Delete
from hbasedriver.operations.get import Get
from hbasedriver.operations.put import Put
from hbasedriver.operations.scan import Scan

config = {
    HConstants.ZOOKEEPER_QUORUM: "127.0.0.1:2181",
}

with Client(config) as client:
    with client.get_table(b"default", b"mytable") as table:
        # Write one column
        table.put(Put(b"row1").add_column(b"cf", b"name", b"Alice"))

        # Read one row
        row = table.get(Get(b"row1"))
        if row is not None:
            print(row.get(b"cf", b"name"))

        # Scan a row-key range
        scan = Scan(start_row=b"row1", end_row=b"row9")
        with table.scan(scan) as scanner:
            for result in scanner:
                print(result.rowkey, result.kv)

        # Delete the whole row
        table.delete(Delete(b"row1"))
```

HBase row keys, column families, qualifiers, and values are byte strings. Encode text before writing it and decode it after reading it.

## Configuration

```python
from hbasedriver.client.client import Client
from hbasedriver.hbase_constants import HConstants

config = {
    # Separate multiple ZooKeeper addresses with commas
    HConstants.ZOOKEEPER_QUORUM: "zk1:2181,zk2:2181,zk3:2181",
    HConstants.CONNECTION_POOL_SIZE: 20,
    HConstants.CONNECTION_IDLE_TIMEOUT: 600,
}

client = Client(config)
```

| Key | Required | Default | Description |
| --- | --- | --- | --- |
| `hbase.zookeeper.quorum` | Yes | None | Comma-separated ZooKeeper addresses |
| `hbase.connection.pool.size` | No | `10` | RegionServer pool size setting (not a global socket cap) |
| `hbase.connection.idle.timeout` | No | `300` | Idle connection timeout in seconds |

Use context managers (or `close()` in `finally`) to request cleanup. Some low-level cleanup/cache paths are incomplete; do not infer full socket reclamation, thread safety, or automatic failover. See [configuration and lifecycle](https://innovationb1ue.github.io/hbase-driver/configuration/).

## Common operations

The snippets below assume an open `table` with a `cf` family. They write demonstration data; use a disposable table. See the [runnable cookbook](https://innovationb1ue.github.io/hbase-driver/cookbook/) for independent examples with assertions.

### Put and Get

```python
from hbasedriver.operations.get import Get
from hbasedriver.operations.put import Put

put = (
    Put(b"user:1001")
    .add_column(b"cf", b"name", b"Alice")
    .add_column(b"cf", b"age", b"30")
)
table.put(put)

get = (
    Get(b"user:1001")
    .add_column(b"cf", b"name")
    .add_column(b"cf", b"age")
)
row = table.get(get)
if row is not None:
    print(row.get(b"cf", b"name"))
```

### Scan and filters

```python
from hbasedriver.filter import PrefixFilter
from hbasedriver.operations.scan import Scan

scan = (
    Scan(start_row=b"user:")
    .with_end_row(b"user;", inclusive=False)
    .add_family(b"cf")
    .set_filter(PrefixFilter(b"user:"))
)

with table.scan(scan) as scanner:
    for row in scanner.next_batch(100):
        print(row.rowkey)
```

Other server-side filters include `RowFilter`, `FamilyFilter`, `QualifierFilter`, `ValueFilter`, `PageFilter`, `ColumnPrefixFilter`, and `SingleColumnValueFilter`.

### Batch operations

```python
from hbasedriver.operations.batch import BatchGet, BatchPut
from hbasedriver.operations.put import Put

batch_put = BatchPut()
batch_put.add_put(Put(b"row1").add_column(b"cf", b"q", b"value1"))
batch_put.add_put(Put(b"row2").add_column(b"cf", b"q", b"value2"))
write_results = table.batch_put(batch_put)

batch_get = BatchGet([b"row1", b"row2"])
batch_get.add_column(b"cf", b"q")
rows = table.batch_get(batch_get)
```

BatchGet returns a dictionary keyed by row key; `add_column()` is not chainable. Batch operations currently issue per-row requests, not one combined RPC.

Use `BufferedMutator` to defer writes in memory, with an explicitly owned `table`:

```python
from hbasedriver.client.buffered_mutator import BufferedMutator, BufferedMutatorParams
from hbasedriver.operations.put import Put

params = BufferedMutatorParams()
params.write_buffer_periodic_flush = 0
with BufferedMutator(table, params) as mutator:
    for index in range(100):
        mutator.mutate(
            Put(f"buffer:{index:04d}".encode("ascii")).add_column(b"cf", b"q", b"value")
        )
    mutator.flush()
```

Flush/close can fail after partial writes; the buffer is not a durable retry queue and does not currently reduce the number of mutation RPCs.

### Atomic operations

```python
from hbasedriver.operations.increment import CheckAndPut, Increment
from hbasedriver.operations.put import Put

increment = Increment(b"counter-row").add_column(b"cf", b"count", 1)
new_value = table.increment(increment)

check_and_put = CheckAndPut(b"lock-row")
check_and_put.set_check(b"cf", b"lock", None)  # Cell must be absent.
check_and_put.set_put(
    Put(b"lock-row").add_column(b"cf", b"lock", b"locked")
)
success = table.check_and_put(check_and_put)
```

The driver also exposes `Append`, `table.check_and_delete()`, `RowMutations`, `exists`, and `exists_all`. Initialize counters from absent cells, not text `b"0"`; replaying Increment/Append after an unknown outcome can duplicate effects.

## Table administration

Use a unique disposable table for a create/delete demonstration. This example only removes a table after its own creation succeeded:

```python
from uuid import uuid4

from hbasedriver.client.client import Client
from hbasedriver.common.table_name import TableName
from hbasedriver.hbase_constants import HConstants
from hbasedriver.operations.column_family_builder import ColumnFamilyDescriptorBuilder

config = {HConstants.ZOOKEEPER_QUORUM: "127.0.0.1:2181"}
table_name = TableName.value_of(b"default", f"docs_admin_{uuid4().hex}".encode("ascii"))
created = False
with Client(config) as client:
    with client.get_admin() as admin:
        try:
            admin.create_table(table_name, [ColumnFamilyDescriptorBuilder(b"cf").build()])
            created = True
            print(admin.describe_table(table_name))
        finally:
            if created:
                admin.disable_table(table_name)
                admin.delete_table(table_name)
```

If creation times out, inspect the unique table name before cleanup: the server may have created it. For namespaces, column families, snapshots, and Region operations, read the [administration guide](https://innovationb1ue.github.io/hbase-driver/admin_guide/). Procedure IDs do not prove completion; truncate does not preserve full table metadata or splits.

## Complete example

The [complete example tutorial](https://innovationb1ue.github.io/hbase-driver/complete_example/) contains a self-contained script with temporary-table cleanup. Copy it to a local file named `docs_quickstart.py`, then run against a development cluster:

```bash
HBASE_ZK=127.0.0.1:2181 python docs_quickstart.py
```

The tutorial verifies schema inspection, CRUD, filters, batches, counters, conditional writes, and buffered writes. JSON, pagination, Append, and RowMutations have additional [cookbook recipes](https://innovationb1ue.github.io/hbase-driver/cookbook/). This documentation-only publication does not update the legacy Python example in the repository root.

## Development and testing

The repository includes a Docker Compose environment with ZooKeeper, an HBase Master, three RegionServers, and a Python test container. Docker and Docker Compose are required.

Start or reuse this project's containers and run the existing related integration tests:

```bash
docker compose up -d
docker compose exec -T dev python -m pytest test/test_batch_operations.py test/test_atomic_operations.py test/test_buffered_mutator.py -v
```

The `dev` service is the `hbase_dev` container; HBASE_ZK is already `hbase-zk:2181`. Integration tests run there because the cluster advertises Docker-network hostnames.

For the full suite and readiness checks on the existing cluster:

```bash
./scripts/run_tests_3node.sh --no-start
```

Without `--no-start`, the runner rebuilds/restarts the project containers. To stop without deleting data, use `docker compose stop`; the runner's `--down` flag removes volumes.

## Documentation

The [official documentation website](https://innovationb1ue.github.io/hbase-driver/) provides navigation, English/Chinese entry points, search, copyable examples, and light/dark appearance. It updates from main independently of PyPI releases.

Start with the [documentation index](https://innovationb1ue.github.io/hbase-driver/) or [中文使用手册](https://innovationb1ue.github.io/hbase-driver/中文介绍/).

- [Getting started](https://innovationb1ue.github.io/hbase-driver/getting_started/) · [Data model and encoding](https://innovationb1ue.github.io/hbase-driver/data_model/)
- [Configuration and lifecycle](https://innovationb1ue.github.io/hbase-driver/configuration/) · [Runnable recipes](https://innovationb1ue.github.io/hbase-driver/cookbook/)
- [Scanning and advanced writes](https://innovationb1ue.github.io/hbase-driver/advanced_usage/) · [API reference](https://innovationb1ue.github.io/hbase-driver/api_reference/)
- [Administration](https://innovationb1ue.github.io/hbase-driver/admin_guide/) · [Current limitations](https://innovationb1ue.github.io/hbase-driver/limitations/)
- [Performance](https://innovationb1ue.github.io/hbase-driver/performance_guide/) · [Troubleshooting](https://innovationb1ue.github.io/hbase-driver/troubleshooting/)
- [Java/Thrift migration](https://innovationb1ue.github.io/hbase-driver/migration_guide/) · [HappyBase migration](https://innovationb1ue.github.io/hbase-driver/migration_from_happybase/)
- [Development environment](https://innovationb1ue.github.io/hbase-driver/DEV_ENV/) · [Test guide](https://innovationb1ue.github.io/hbase-driver/TEST_GUIDE/) · [Documentation maintenance](https://innovationb1ue.github.io/hbase-driver/documentation_guide/)

Current limitations include Scan limit/time-range handling, Scan column serialization, flattened version results, and incomplete secure/HA and resource-sharing behavior. Passing the cookbook does not certify all API options or production topologies. These docs track repository code; PyPI metadata changes take effect in a future release.

## License

This project is licensed under the [Apache License 2.0](https://github.com/innovationb1ue/hbase-driver/blob/main/LICENSE).
