Metadata-Version: 2.4
Name: gigaxml
Version: 1.1.0
Summary: A production-oriented CLI toolkit for profiling, validating and extracting structured data from multi-gigabyte XML files with bounded memory usage.
License-Expression: MIT
Project-URL: Homepage, https://github.com/gg320324492-lgtm/GigaXML-Memory-Efficient-XML-Extraction-Toolkit
Keywords: xml,streaming,etl,iterparse,large-files,cli,memory-efficient
Classifier: Development Status :: 5 - Production/Stable
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Text Processing :: Markup :: XML
Classifier: Typing :: Typed
Requires-Python: <3.14,>=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: lxml>=5.0
Requires-Dist: pyyaml>=6.0
Provides-Extra: parquet
Requires-Dist: pyarrow>=14.0; extra == "parquet"
Provides-Extra: pretty
Requires-Dist: rich>=13.0; extra == "pretty"
Provides-Extra: gui
Requires-Dist: PySide6>=6.6; extra == "gui"
Requires-Dist: pytest-qt>=4.4; extra == "gui"
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: pytest-cov>=5.0; extra == "dev"
Requires-Dist: ruff>=0.6; extra == "dev"
Requires-Dist: psutil>=5.9; extra == "dev"
Requires-Dist: pyarrow>=14.0; extra == "dev"
Requires-Dist: rich>=13.0; extra == "dev"
Requires-Dist: pyinstaller>=6.10; extra == "dev"
Requires-Dist: pillow>=10.0; extra == "dev"
Dynamic: license-file

# GigaXML

A production-oriented CLI toolkit for profiling, validating and extracting structured data
from multi-gigabyte XML files with bounded memory usage.

## What this is

The point of this project is not "it can parse XML" — plenty of tools can. The point is
**constant, bounded peak memory while extracting from 4 GB / 10 GB files**, and being able
to prove it with a reproducible measurement harness.

Measured on this machine, on a generated 4.05 GiB file holding 11,915,264 records:

| | |
|---|---|
| **Input** | 4142.72 MiB, 11,915,264 records |
| **Time** | 284.67 s (14.6 MiB/s, 41,856 records/s) |
| **Peak RSS** | 33.703 MiB |
| **Increase over the post-import baseline** | **5.059 MiB** |

The same run at 1 GiB (2,978,816 records) added **5.105 MiB** — four times the input and
the increase did not move. Something that accumulated per record would make the four
gigabyte figure four times the one gigabyte figure; it is flat to within a percent. Every
number here comes from a script in this repository; see [Benchmarks](#benchmarks).

The configuration used above reads six fields, including a nested path and a type
conversion — the shapes a real config uses:

```yaml
record: /catalog/products/product
fields:
  id:           {path: '@id'}
  type:         {path: '@type'}
  name:         {path: name}
  category:     {path: category}
  price:        {path: price, type: float}
  manufacturer: {path: manufacturer/name}
```

## Install

From PyPI:

```bash
pip install gigaxml           # the command-line toolkit
pip install "gigaxml[gui]"    # and the desktop application (pulls PySide6: 640 MiB installed, measured on Windows)
```

Parquet output needs the `parquet` extra (`pip install "gigaxml[parquet]"`); CSV and JSONL
do not. The desktop application is also packaged per platform — an unsigned Windows
build, an Apple Silicon `.dmg` and an x86_64 AppImage — under
[Releases](https://github.com/gg320324492-lgtm/GigaXML-Memory-Efficient-XML-Extraction-Toolkit/releases);
its release notes say what each build runs on and what the unsigned warnings mean.

From source, if you would rather:

```bash
git clone https://github.com/gg320324492-lgtm/GigaXML-Memory-Efficient-XML-Extraction-Toolkit.git
cd GigaXML-Memory-Efficient-XML-Extraction-Toolkit
python -m venv .venv
.venv/Scripts/python -m pip install -e ".[dev]"     # POSIX: .venv/bin/python
```

`lxml` and `pyyaml` are the only required dependencies.

## Thirty seconds

Point it at a document you know nothing about, let it propose a config, then run it:

```bash
# 1. What is in this file?
gigaxml inspect big.xml

# 2. Write a starting-point config for the highest-ranked record candidate
gigaxml inspect big.xml --generate-config config.yaml --infer-types

# 3. Extract
gigaxml extract big.xml -c config.yaml -o out.csv
```

`inspect` does not read the whole document — it reports the structure and the repeating
paths it found, and ranks them. The config it writes is explicitly a **starting point**,
not a conclusion; read the comments in it.

For a quick look at the data before committing to a full run:

```bash
gigaxml sample big.xml -c config.yaml -n 20 -o first20.jsonl
```

### What it looks like

![gigaxml inspect reading the structure of a document](assets/inspect.gif)

*Finding the records in a document that opens with a licence comment.*

![gigaxml extract processing a four-gigabyte file](assets/extract-4g.gif)

*Extracting 11.9 million records from 4.05 GiB.*

> **Both of these are animations rendered from the tools' real output, not screen
> recordings.** The text is what the tools actually printed and the timings are the
> measured ones, but the frames are drawn rather than captured — this machine's sandbox
> does not permit screen capture. Each frame carries the same note.

## Benchmarks

Three generated datasets, two configs, measured in a subprocess with `psutil`:

```
dataset   fields     input MiB      records        s   MiB/s     peak    delta
------------------------------------------------------------------------------
100MB     6 fields      100.57      290,900     6.92    14.5   33.281    5.113
100MB     1 field       100.57      290,900     2.69    37.4   31.203    2.617
1GB       6 fields     1033.65    2,978,816    70.66    14.6   33.676    5.105
1GB       1 field      1033.65    2,978,816    27.62    37.4   31.496    2.922
4GB       6 fields     4142.72   11,915,264   284.67    14.6   33.703    5.059
4GB       1 field      4142.72   11,915,264   112.04    37.0   31.500    2.973
```

`peak` and `delta` are MiB; `delta` is against the same process's post-import baseline.
`peak` is `PeakWorkingSetSize` — the maximum over the process's life, not the current RSS
at the end, which reads 10–14% lower.
The one-field rows are a control, not the headline: a single-field config is the easiest
member of this family to run, and quoting it alone would overstate what a real config
costs. Field count costs about **2.5×** in throughput.

To reproduce, generate the datasets and run the harness in [`benchmarks/`](benchmarks):

```bash
gigaxml generate --size 100MB -o data/b100m.xml
gigaxml generate --size 1GB   -o data/b1g.xml
gigaxml generate --size 4GB   -o data/b4g.xml
python benchmarks/bench_extraction.py
```

`--size` is approximate: `--size 1GB` produces 1033.65 MiB, not 1024, and the sizes above
are the measured ones. Peak and delta are both reported because either alone can be
misread — peak includes about 32 MiB of interpreter and library overhead, delta is what
the workload is responsible for, and both baselines in this repository are taken after
every import so that two deltas are comparable.

## Non-goals

- No full XPath 3.1 — XPath is evaluated only inside a single record subtree.
- No arbitrary byte-offset seek/resume — XML byte offsets are not a safe parse boundary.
- No AI/ML structure inference — confidence values are deterministic statistics.
- No real customer data — everything runs on synthetic, reproducible datasets.
- No fabricated benchmarks — every performance claim comes from a runnable script.

## Known limitations

- **`--resume` re-parses and skips; it does not seek.** XML cannot be re-entered
  mid-stream, so continuing a run means reading from the beginning and discarding the
  records already accounted for. On a 403 MB file that costs 8.7 s against 17.1 s to
  extract, so resuming saves roughly half of what you had already done. `--help` says so
  too.
- **`inspect` is slower than `extract`** — 17.3 MiB/s against 38.2 MiB/s on the same
  1 GB file. It maintains several parallel bookkeeping stacks per element. It is also the
  command you run once on a document, not in a loop.
- **An inferred config treats containers as leaves.** `--generate-config` proposes direct
  children and attributes; a field whose element has children of its own is read as
  concatenated text, so `<tags><tag>a</tag><tag>b</tag></tags>` becomes `ab`. Nested
  paths (`manufacturer/name`) have to be written by hand, as the generated comments say.
- **Types are inferred from a sample**, and `decimal` is never inferred. If a field is
  money, set `type: decimal` yourself — `float` cannot represent 49.90 exactly.
- **Parsing limits are not configurable.** Entities are never expanded, the network is
  never touched, and no DTD is loaded; documents nested deeper than 256 levels, carrying a
  single text node over about 10 MB, or amplified by entities are refused rather than
  partially read. These are deliberate and there are no flags to turn them off.
- **`--checkpoint-every` verifies the parts on disk before resuming**, which costs one
  pass over the output at about **200 MiB/s** — about 10 ms for 2 MiB of CSV, negligible for
  Parquet, whose row counts come from file metadata. It grows with the size of the
  output, not the input. Measured by `benchmarks/bench_resident.py`.
- **A resume is dominated by starting the process, not by checking the output.** Against
  an already-complete manifest on a 403 MiB source, the command takes about 770 ms: some
  400 ms of that is interpreter startup and imports, and most of the rest is hashing the
  source to confirm it has not changed. The part check itself is about 10 ms.
- **One field value that is itself gigabytes is held in memory.** Records stream, but
  there is no streaming mode for a single value, because there is nothing to stream it
  into.
- **`sample` reads what it samples into memory.** It is meant for looking at a file, not
  for measuring one.
- **Input is a local file, never a URL.** A `.xml.gz` file is fine; a document that lives
  behind HTTP is out of scope.

## Development

```bash
pytest -q                                   # unit + integration
pytest -q tests/performance                 # memory and throughput; not in CI
ruff check .
ruff format --check .
```

Performance tests are excluded from CI: they measure memory and throughput, take minutes,
and are not a pass/fail signal.

## License

MIT — see [LICENSE](LICENSE).
