Metadata-Version: 2.4
Name: criba
Version: 0.1.0
Summary: Read real-world spreadsheets without silent corruption: European numbers, ambiguous dates, messy column names.
Author: Bryan Ross
License: MIT License
        
        Copyright (c) 2026 Bryan Ross
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
        
Project-URL: Homepage, https://github.com/Bryanbackoffice01/criba
Project-URL: Source, https://github.com/Bryanbackoffice01/criba
Project-URL: Issues, https://github.com/Bryanbackoffice01/criba/issues
Keywords: pandas,csv,excel,spreadsheet,locale,i18n,decimal-separator,dayfirst,data-cleaning,european
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Office/Business :: Financial :: Accounting
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Classifier: Topic :: Text Processing :: General
Classifier: Typing :: Typed
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pandas>=1.5
Provides-Extra: excel
Requires-Dist: openpyxl>=3.0; extra == "excel"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: openpyxl>=3.0; extra == "dev"
Requires-Dist: build; extra == "dev"
Requires-Dist: twine; extra == "dev"
Dynamic: license-file

# criba

Read real-world spreadsheets without silent corruption.

*[Léeme en español](README.es.md)*

---

## The problem

```python
>>> import pandas as pd
>>> pd.to_datetime(pd.Series(["2026-01-08"]), dayfirst=True)[0]
Timestamp('2026-08-01 00:00:00')
```

That is an ISO date — the least ambiguous format there is — turned into a
different day. No exception. No warning in the result. And `dayfirst=True`
is what you set when your data comes from Spain, Germany, France, Italy,
Brazil or most of the rest of the world.

It is also a regression: **pandas 1.5 read that date correctly, and 2.0
started getting it wrong.** Every version since behaves this way, up to and
including 3.0.

The dates that fail outright become `NaT`, which is worse than it sounds:
rows with `NaT` all look alike, so a `groupby` will happily merge records
that have nothing to do with each other.

Amounts have the same problem from the other direction. `1.234,56` is how a
billion people write "one thousand two hundred thirty-four and fifty-six
cents". `pd.to_numeric` returns `NaN` for it.

## What criba does

Guesses the same things, but from the whole column instead of one value at
a time — and then tells you what it decided.

```python
import criba

table, report = criba.read("facturas.csv", explain=True)
print(report.summary())
```

```
facturas.csv: 312 rows, 5 columns
  encoding cp1252, separator ';'
  dates read as day-first (from the data)
  date='Fecha', amount='Importe', party='Proveedor', reference='Nº Factura'
```

If any of those is wrong, you find out now rather than three steps later.

## Install

```bash
pip install criba            # CSV
pip install criba[excel]     # also .xlsx
```

Only dependency is pandas.

## The parts

### Dates that stay the day they were

```python
>>> criba.to_datetime(pd.Series(["2026-01-08"]))[0]
Timestamp('2026-01-08 00:00:00')
```

The useful trick is that one unambiguous value settles the whole column.
`25/12/2026` can only be day-first — there is no month 25 — so every other
value in that column is read the same way:

```python
>>> dates = criba.to_datetime(pd.Series(["25/12/2026", "08/01/2026"]))
>>> dates[1].month
1
```

When the data cannot settle it, you get a warning rather than a quiet guess.
When the column is written both ways — which happens in files assembled by
hand — you get told that too:

```python
>>> found = criba.detect_date_order(pd.Series(["25/12/2026", "12/25/2026"]))
>>> found.warnings[0]
'mixed date orders: 1 values can only be day-first and 1 can only be month-first; ...'
```

### Numbers as people actually write them

```python
>>> criba.to_number("1.234,56")     # Spain, Germany, Italy, Brazil
1234.56
>>> criba.to_number("1,234.56")     # UK, US
1234.56
>>> criba.to_number("1 234,56")     # France
1234.56
>>> criba.to_number("€ 1.234,56")
1234.56
>>> criba.to_number("(1.234,56)")   # accounting negative
-1234.56
>>> criba.to_number("1'234.56")     # Switzerland
1234.56
>>> criba.to_number("1,23E+05")     # what spreadsheets do to large numbers
123000.0
>>> criba.to_number("12,34,56,78")  # not a number in any locale
nan
```

That last one matters. Anything that is not unambiguously a number becomes
`NaN`, never a plausible-looking wrong value.

### Column names

```python
>>> df = pd.DataFrame(columns=["F. Emisión", "Razón Social", "Importe total (EUR)"])
>>> found = criba.detect_columns(df)
>>> found.amount
'Importe total (EUR)'
```

Spanish, English and Portuguese profiles are included; pass your own mapping
for anything else. Nothing it cannot place is filled in with a guess — it
stays `None`, and `found.missing("date", "reference")` tells you what is
absent.

Whatever you pass explicitly always wins:

```python
criba.read("facturas.csv", amount="TOTAL FRA")
```

And when two columns both fit, you are told which one was used and what the
other one was, because choosing between them quietly is the mistake:

```python
>>> df = pd.DataFrame(columns=["Importe", "Importe total"])
>>> criba.detect_columns(df).why()[0]
"amount: used 'Importe', but 'Importe total' also fitted"
```

### Workbooks that open on a cover page

A spreadsheet whose first sheet holds a title and a logo used to give you a
one-column table called `nota`, with no error. criba uses the sheet that
most looks like a table, and says which:

```python
>>> table, report = criba.read("cuentas.xlsx", explain=True)  # doctest: +SKIP
>>> report.sheet                                              # doctest: +SKIP
'Datos'
```

`sheet="Datos"` or `sheet=1` chooses it yourself.

### Tables that do not start on line 1

Accounting software rarely puts the column names at the top:

```
INFORME DE FACTURACIÓN 2026
Generado el 19/09/2026

Fecha;Proveedor;Importe
25/12/2026;Luis;1.234,56
```

Read straight, that file either fails outright or gives you a one-column
table called `INFORME DE FACTURACIÓN 2026`. criba finds the real header and
says where it was:

```python
>>> table, report = criba.read("informe.csv", explain=True)   # doctest: +SKIP
>>> report.header_row                                         # doctest: +SKIP
4
```

The rule is deliberately hard to fool: the header is the first row as wide
as the body of the table, made of labels rather than values, with no
repeated names, and followed by actual data. Pass `header_row=4` to say
where it is instead of letting it look.

### Row numbers your sender can find

Every table gets a `_source_row` column holding the row number **as it
appears in their spreadsheet** — header counted, starting at 1. Telling
somebody "row 47" is only useful if row 47 is what they see when they open
the file.

## Scope

criba reads files. It does not clean, validate or analyse them — there are
good libraries for that, and this one is meant to be the thing you run
before them.

## Licence

MIT.
