Metadata-Version: 2.4
Name: gffbase
Version: 0.3.0
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Programming Language :: Rust
Classifier: Topic :: Scientific/Engineering
Classifier: Topic :: Scientific/Engineering :: Bio-Informatics
Classifier: Topic :: Database
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Dist: duckdb>=1.4.1
Requires-Dist: pyarrow>=18.1
Requires-Dist: pandas>=1.5 ; extra == 'all'
Requires-Dist: polars>=0.20 ; extra == 'all'
Requires-Dist: pyfaidx>=0.7 ; extra == 'all'
Requires-Dist: pybedtools>=0.9 ; sys_platform != 'win32' and extra == 'all'
Requires-Dist: biopython>=1.80 ; extra == 'all'
Requires-Dist: matplotlib>=3.5 ; extra == 'all'
Requires-Dist: psutil>=5 ; extra == 'bench'
Requires-Dist: memory-profiler>=0.61 ; extra == 'bench'
Requires-Dist: gffutils>=0.13 ; extra == 'bench'
Requires-Dist: biopython>=1.80 ; extra == 'biopython'
Requires-Dist: pytest>=7 ; extra == 'dev'
Requires-Dist: pytest-cov>=4 ; extra == 'dev'
Requires-Dist: hypothesis>=6 ; extra == 'dev'
Requires-Dist: packaging>=24 ; extra == 'dev'
Requires-Dist: psutil>=5 ; extra == 'dev'
Requires-Dist: ruff>=0.4 ; extra == 'dev'
Requires-Dist: mypy>=1.8 ; extra == 'dev'
Requires-Dist: maturin>=1.5,<2.0 ; extra == 'dev'
Requires-Dist: sphinx==9.1.0 ; extra == 'docs'
Requires-Dist: furo==2025.12.19 ; extra == 'docs'
Requires-Dist: sphinx-copybutton==0.5.2 ; extra == 'docs'
Requires-Dist: sphinx-design==0.7.0 ; extra == 'docs'
Requires-Dist: pyfaidx>=0.7 ; extra == 'fasta'
Requires-Dist: pandas>=1.5 ; extra == 'pandas'
Requires-Dist: pybedtools>=0.9 ; sys_platform != 'win32' and extra == 'plot'
Requires-Dist: matplotlib>=3.5 ; extra == 'plot'
Requires-Dist: polars>=0.20 ; extra == 'polars'
Requires-Dist: pybedtools>=0.9 ; sys_platform != 'win32' and extra == 'pybedtools'
Requires-Dist: pytest>=7 ; extra == 'test'
Requires-Dist: pytest-cov>=4 ; extra == 'test'
Requires-Dist: hypothesis>=6 ; extra == 'test'
Requires-Dist: packaging>=24 ; extra == 'test'
Requires-Dist: pyyaml>=6 ; extra == 'test'
Requires-Dist: psutil>=5 ; extra == 'test'
Requires-Dist: tomli>=2 ; python_full_version < '3.11' and extra == 'test'
Provides-Extra: all
Provides-Extra: bench
Provides-Extra: biopython
Provides-Extra: dev
Provides-Extra: docs
Provides-Extra: fasta
Provides-Extra: pandas
Provides-Extra: plot
Provides-Extra: polars
Provides-Extra: pybedtools
Provides-Extra: test
License-File: LICENSE
Summary: GFFBase — Rust-accelerated GFF3/GTF parser with a DuckDB-backed storage engine and a drop-in gffutils-compatible Python API.
Keywords: gff,gff3,gtf,gencode,bioinformatics,genomics,annotation,duckdb,rust,pyo3,parser,feature-database
Home-Page: https://github.com/Kuanhao-Chao/gffbase
Author-email: Kuan-Hao Chao <kuanhao.chao@gmail.com>
License: Apache-2.0
Requires-Python: >=3.10
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Changelog, https://github.com/Kuanhao-Chao/gffbase/releases
Project-URL: Documentation, https://khchao.com/gffbase/
Project-URL: Homepage, https://github.com/Kuanhao-Chao/gffbase
Project-URL: Issues, https://github.com/Kuanhao-Chao/gffbase/issues
Project-URL: Source, https://github.com/Kuanhao-Chao/gffbase

<p align="center">
  <img src="https://raw.githubusercontent.com/Kuanhao-Chao/gffbase/main/docs/source/_static/logo.svg#gh-light-mode-only" alt="gffbase" width="62%">
  <img src="https://raw.githubusercontent.com/Kuanhao-Chao/gffbase/main/docs/source/_static/logo-white.svg#gh-dark-mode-only" alt="gffbase" width="62%">
</p>


[![PyPI version](https://img.shields.io/pypi/v/gffbase.svg)](https://pypi.org/project/gffbase/)
[![PyPI downloads](https://img.shields.io/pypi/dm/gffbase.svg)](https://pypi.org/project/gffbase/)
[![Python versions](https://img.shields.io/pypi/pyversions/gffbase.svg)](https://pypi.org/project/gffbase/)
[![License](https://img.shields.io/badge/license-Apache%202.0-blue.svg)](https://github.com/Kuanhao-Chao/gffbase/blob/main/LICENSE)
[![Docs](https://img.shields.io/badge/docs-khchao.com%2Fgffbase-blue.svg)](https://khchao.com/gffbase/)
[![CI](https://github.com/Kuanhao-Chao/gffbase/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/Kuanhao-Chao/gffbase/actions/workflows/ci.yml)

## What is GFFBase?

**A high-performance genomic-annotation engine — a SIMD Rust parser, a DuckDB
columnar backend and a zero-copy PyArrow interface — built for whole-genome
ingest and bulk machine-learning feature extraction, and a drop-in successor to
[`gffutils`](https://github.com/daler/gffutils).**

The legacy `FeatureDB` / `Feature` / `create_db` / `DataIterator` / `GFFWriter` /
`merge_criteria` API is preserved verbatim, so most scripts migrate by changing
one import line.

## Install

```bash
pip install gffbase
```

Universal `abi3` wheels: one binary per platform covers CPython 3.10 through
3.14, and no Rust toolchain is needed. Source and development builds are in the
[installation guide](https://khchao.com/gffbase/content/installation.html).

> **0.3.0 is the current release** — ingest 2.4–2.9× faster, in under a third
> of the memory on whole-genome files; loops over a live stream prefetched; and
> noisy files read as they were meant. Coming from 0.2.0, upgrade: 0.2.1 fixed
> defects that silently lost data. Coming from 0.1.0, 0.2 and later fix two SQL
> injection vulnerabilities ([advisory](https://github.com/Kuanhao-Chao/gffbase/security/advisories/GHSA-5f5g-g3v5-prrg))
> and have breaking changes. All are in the
> [changelog](https://github.com/Kuanhao-Chao/gffbase/blob/main/CHANGELOG.md).

## Quick start

<!-- docs-test: skip reason="needs the GENCODE v49 corpus" -->
```python
from gffbase import create_db

# Ingest a GTF or GFF3 — format auto-detected, gzip transparent.
with create_db("gencode.v49.basic.annotation.gtf.gz", "gencode.duckdb") as db:

    # Walk one gene's hierarchy. GENCODE ids carry their version.
    for tx in db.children("ENSG00000139618.19", level=1, featuretype="transcript"):
        print(tx.id, tx.start, tx.end)

    # Overlap query, routed to the per-seqid R-tree.
    for exon in db.region("chr17:43044295-43125483", featuretype="exon"):
        print(exon)
```

Coming from `gffutils`? `import gffbase as gffutils` is usually the only change
you need — but read the [migration guide](https://khchao.com/gffbase/content/migration.html) first,
for the one loop pattern you should not port unchanged.

## Bulk extraction without Python objects

The workload GFFBase exists for: pull every exon for tens of thousands of
transcripts, hand the columns to a tensor, train. One set-based query returns
DuckDB's own Arrow buffers, and **no** `Feature` **object is constructed at any
layer** — where a per-feature Python loop allocates millions of throwaway
objects and that allocation dominates everything else.

<!-- docs-test: skip reason="illustrative: an id list the reader supplies" -->
```python
exons = db.children_batched(
    transcript_ids,              # an iterable of 50 000 IDs
    featuretype="exon",
    format="arrow",              # "df" and "polars" also supported
)
# A pyarrow.Table sharing memory with DuckDB. Its "anchor" column carries the
# input id for each row, so per-transcript groups survive without N queries.

import torch
starts = torch.from_numpy(exons.column("start").to_numpy())
```

`region_batched()` and `parents_batched()` offer the same zero-copy contract for
spatial and parent workloads. One trade to know: a loop over a gffbase iterator
(`for g in db.features_of_type("gene"): db.children(g)`) is prefetched and runs
within 1.4–3.5× of `gffutils`, but a loop over a list of ids of your own,
`for i in ids: db.children(i)`, is **slower** — each call is a DuckDB statement
of its own — which is why the batched API exists.

End-to-end PyTorch and Hugging Face pipelines are in the
[ML workflows cookbook](https://khchao.com/gffbase/content/cookbook_ml_workflows.html),
and every method has a snippet in the
[usage gallery](https://khchao.com/gffbase/content/usage_gallery.html).

## Measured against `gffutils`

Head-to-head across five canonical human-genome annotations — one run, one
machine, one commit. **gffbase ingests every corpus faster, 1.92× to 3.62×,**
into a database 0.61× to 0.89× the size of gffutils' SQLite file. The narrowest
margin is the GENCODE GTF row, the arm least favourable to gffbase: with parent
inference off, gffutils' GTF path is a plain bulk insert. Batched extraction,
spatial indexing and SQL over the whole corpus come on top.

<!-- BEGIN GENERATED: corpus-table -->
| Corpus | Format | Lines | gffbase ingest | legacy ingest | speedup | peak RSS (ingest + full validation) | spatial qps | batched (5 k anchors) |
| --- | :--: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| **GENCODE v49** (basic) | GTF | 6,068,892 | **3 min 36 s** | 6 min 54 s | **1.92×** | 53.47 GB | **728** ±1% (n=5) | 614 ms / 1.93 M desc |
| **GENCODE v49** (basic) | GFF3 | 6,066,054 | **3 min 13 s** | 9 min 45 s | **3.03×** | 61.93 GB | **770** ±1% (n=5) | 666 ms / 1.93 M desc |
| **RefSeq GRCh38.p14** | GFF3 | 4,932,571 | **2 min 5 s** | 6 min 31 s | **3.13×** | 26.69 GB | **588** ±1% (n=5) | 443 ms / 999 k desc |
| **CHESS 3.1.3** | GFF3 | 2,761,061 | **37.0 s** | 2 min 14 s | **3.62×** | 2.93 GB | **678** ±3% (n=5) | 155 ms / 161 k desc |
| **MANE v1.5** (Ensembl) | GFF3 | 524,834 | **15.6 s** | 45.1 s | **2.89×** | 4.00 GB | **890** ±0% (n=5) | 136 ms / 156 k desc |
<!-- END GENERATED: corpus-table -->

<!-- BEGIN GENERATED: benchmark-provenance -->
**Measured on** AMD EPYC 7702 64-Core Processor · 128 cores · 1007.22 GB RAM · Linux-5.14.0-503.15.1.el9_5.x86_64-x86_64-with-glibc2.34  
**Versions:** Python 3.11.16 · gffbase 0.3.0 · duckdb 1.5.5 · pyarrow 25.0.1 · gffutils 0.14  
**Commit:** `632a4d80dee0` · **Run:** 2026-09-27T15:07:53Z  
*Generated from `benchmarks/results/06_mega.linux-x86_64.json` by `tools/gen_benchmark_tables.py`. Do not edit by hand.*
<!-- END GENERATED: benchmark-provenance -->

`peak RSS` is ingest **plus exhaustive validation**, not the cost of ingest:
`validate_db` defaults to `sample=200` and the CLI never overrides it. No
speedup is published unless both engines' correctness signatures match first,
and every number is generated from a committed measurement file — a test fails
if a published table stops matching it. The per-corpus analysis, the fairness
constraints and the measured run-to-run spread are on the
[performance](https://khchao.com/gffbase/content/performance.html) and
[methodology](https://khchao.com/gffbase/content/methodology.html) pages.

## What's inside

- **Rust + PyO3 parser** — streaming, SIMD line and tab splitting,
  percent-decoding in and encoding back out, GTF semicolon-in-quotes safe, gzip
  and one-file tar archives transparent, failures raised as line-numbered
  `GFFFormatError`.
- **Rust ingest producer** — ids, duplicates and batching without a Python
  object per record, straight into DuckDB as Arrow; memory bounded by a budget.
- **DuckDB columnar storage** — an 11-table schema plus 3 compatibility views,
  set-based GTF gene/transcript synthesis, a materialized transitive closure,
  and a per-seqid-banded R-tree built inline during ingest.
- **Smart routing** — `region()` picks the R-tree or a zone-map-pruned scan;
  `children()` and `parents()` look ids up in one table, and inside a loop over
  a gffbase iterator are prefetched.
- **Noisy input** — byte-order marks, CR line endings, CDS-only and AUGUSTUS
  GTF, Liftoff ids, truncated gzip: read as meant, and reported
  ([noisy files](https://khchao.com/gffbase/content/noisy_files.html)).
- **Vectorized batched API** — `children_batched`, `parents_batched` and
  `region_batched` return `pyarrow.Table`, `pandas.DataFrame` or
  `polars.DataFrame` straight out of DuckDB's buffer pool.
- **Drop-in legacy API** — plus `bed12`, `interfeatures`, an `execute()` SQL
  escape hatch and `export_sqlite()`, with every remaining difference from
  `gffutils` 0.14 declared and tested.
- **Discontinuous features** — several lines sharing one `ID` are one logical
  feature, with per-segment phase, `covered_length` and `explode_segments=`.
- **A command line and validation** — `gffbase create|fetch|children|parents|region|search|rmdups|sanitize|validate|migrate`,
  structural validation at two depths, and in-place migration of a v1 database.

## Documentation

Full site: **[khchao.com/gffbase](https://khchao.com/gffbase/)**

| Page | What's there |
| --- | --- |
| [Quickstart](https://khchao.com/gffbase/content/quickstart.html) | A ten-minute walkthrough on a demo annotation |
| [Migration guide](https://khchao.com/gffbase/content/migration.html) | Drop-in checklist, and the one loop pattern to batch |
| [Noisy files](https://khchao.com/gffbase/content/noisy_files.html) | What gffbase does with each non-standard input |
| [Tuning](https://khchao.com/gffbase/content/tuning.html) | Threads, the memory budget, ingest stages, when to batch |
| [Usage gallery](https://khchao.com/gffbase/content/usage_gallery.html) | Every public method, copy-pasteable |
| [Cookbooks](https://khchao.com/gffbase/content/cookbooks.html) | GENCODE/Ensembl, RefSeq, MANE, ML workflows |
| [Performance](https://khchao.com/gffbase/content/performance.html) | The numbers, and what they do not claim |
| [Command line](https://khchao.com/gffbase/content/cli.html) | Every `gffbase` subcommand |
| [API reference](https://khchao.com/gffbase/content/api.html) | Full signatures and docstrings |
| [FAQ](https://khchao.com/gffbase/content/faq.html) and [troubleshooting](https://khchao.com/gffbase/content/troubleshooting.html) | Short answers; errors organized by message |

## Contributing

Pull requests, bug reports and feature suggestions are welcome.
[`CONTRIBUTING.md`](https://github.com/Kuanhao-Chao/gffbase/blob/main/CONTRIBUTING.md) covers the Rust and Python
development setup, the test suite and its coverage gates, branch naming and the
PR checklist. The repository ships
[issue](https://github.com/Kuanhao-Chao/gffbase/tree/main/.github/ISSUE_TEMPLATE) and
[pull-request](https://github.com/Kuanhao-Chao/gffbase/blob/main/.github/PULL_REQUEST_TEMPLATE.md) templates.

## License

Apache License 2.0 — see [`LICENSE`](https://github.com/Kuanhao-Chao/gffbase/blob/main/LICENSE).

## Citation

If GFFBase contributes to your research, please cite it:

```bibtex
@software{chao_gffbase_2026,
  author  = {Chao, Kuan-Hao},
  title   = {{GFFBase}: Rust-accelerated GFF3/GTF parser with a
             DuckDB-backed storage engine and zero-copy PyArrow interface},
  year    = 2026,
  version = {0.3.0},
  url     = {https://github.com/Kuanhao-Chao/gffbase},
}
```

The repository also ships a [`CITATION.cff`](https://github.com/Kuanhao-Chao/gffbase/blob/main/CITATION.cff), so
GitHub's "Cite this repository" button produces an up-to-date reference.
Per-version DOIs, when available, are tracked on the [releases page](https://github.com/Kuanhao-Chao/gffbase/releases).

