Metadata-Version: 2.5
Name: docmore-utility
Version: 0.1.0
Summary: 文档与图片辅助工具库：PDF/图片互转、图片/base64 互转、二进制流文件类型识别
Author: mps.wen
License: MIT
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.10
Requires-Dist: pillow>=10.0
Requires-Dist: pymupdf>=1.24
Provides-Extra: magika
Requires-Dist: magika>=1.0.3; extra == 'magika'
Provides-Extra: test
Requires-Dist: magika>=1.0.3; extra == 'test'
Description-Content-Type: text/markdown

# docmore-utility

文档与图片辅助工具库：PDF/图片互转、图片/base64 互转、二进制流文件类型识别。

## 安装

```bash
pip install docmore-utility            # 基础功能（PDF/图片/base64）
pip install docmore-utility[magika]    # 加上文件类型识别
```

## 功能

### PDF ↔ 图片

```python
from docmore_utility import pdf_to_images, images_to_pdf

# PDF 转图片：返回每页 PNG 的字节流
pages = pdf_to_images("doc.pdf", dpi=150, fmt="png")

# 或直接写文件，返回生成的文件路径
paths = pdf_to_images("doc.pdf", output_dir="out/")

# 只渲染指定页（0 起始页索引）、直接拿 PIL 图片、短边不足自动放大、烘焙批注/表单
images = pdf_to_images(
    "doc.pdf",
    pages=[0, 2, 3],
    dpi=150,
    min_dim=1024,      # 每页渲染结果短边至少 1024px（不足自动提高 dpi）
    flatten=True,      # 先把批注/表单字段烘焙进页面再渲染
    return_pil=True,   # 直接返回 list[PIL.Image.Image]（RGB）
)

# 多张图片合并为一个 PDF
images_to_pdf(["1.png", "2.jpg"], output_path="merged.pdf")
```

### 图片 ↔ base64

```python
from docmore_utility import image_to_base64, base64_to_image

b64 = image_to_base64("photo.jpg")                     # 纯 base64 字符串
b64 = image_to_base64("photo.jpg", with_data_uri=True) # data:image/jpeg;base64,...

# 内存中的 PIL 图片（无需落盘），可直接生成 OpenAI 兼容的 data URI
b64 = pil_to_base64(pil_image, format="PNG", with_data_uri=True)

img_bytes = base64_to_image(b64)                 # 返回字节流
path = base64_to_image(b64, output_path="out.png")  # 写文件并返回路径
```

### 文件类型识别（需安装 `[magika]` extra）

```python
from docmore_utility import identify

with open("mystery.bin", "rb") as f:
    result = identify(f.read())

result.label    # "pdf"   -- Magika 内容类型标签
result.mime     # "application/pdf"
result.score    # 识别置信度
result.is_text  # 是否为文本类文件
```

## 开发

```bash
uv sync            # 安装依赖（含 dev 组）
uv run pytest      # 运行测试
uv build           # 构建 wheel / sdist
```

## License

MIT
