Metadata-Version: 2.5
Name: promptflame-optimizer
Version: 0.1.0
Summary: Companion to prompt-flamegraph: automatically trim, deduplicate and minify LLM prompts to fit token budgets
Project-URL: Homepage, https://github.com/fjjjuv/prompt-optimizer
Project-URL: Repository, https://github.com/fjjjuv/prompt-optimizer
Author: Fjjjuv
License: LGPL-3.0-or-later
License-File: LICENSE
License-File: LICENSE.GPL
Keywords: budget,context,llm,optimizer,prompt,token
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: GNU Lesser General Public License v3 or later (LGPLv3+)
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: pytest>=7.0; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Provides-Extra: flamegraph
Requires-Dist: prompt-flamegraph>=0.3.3; extra == 'flamegraph'
Description-Content-Type: text/markdown

<div align="center">

# ✂️ prompt-optimizer

**Cut your LLM prompts down to size — automatically.**

Whitespace cleanup · dedup · tool minification · history & RAG trimming — all local, zero-dependency.

[![PyPI version](https://img.shields.io/pypi/v/promptflame-optimizer)](https://pypi.org/project/promptflame-optimizer/)
[![Python versions](https://img.shields.io/pypi/pyversions/promptflame-optimizer)](https://pypi.org/project/promptflame-optimizer/)
[![License: LGPL v3](https://img.shields.io/badge/license-LGPLv3-blue.svg)](LICENSE)

</div>

[prompt-flamegraph](https://github.com/fjjjuv/prompt-flamegraph) shows you **where** your prompt tokens go. **prompt-optimizer cuts them.** Same structured-prompt format, same zero-dependency philosophy — feed it a prompt, get a smaller prompt plus a report of every change.

> 📦 **PyPI note:** the name `prompt-optimizer` was already taken on PyPI, so this package ships as **`promptflame-optimizer`**. The import name is still `prompt_optimizer` and the CLI is still `prompt-optimizer`.

---

## Install

```bash
pip install promptflame-optimizer
```

| Extra | Why |
|---|---|
| `pip install promptflame-optimizer[flamegraph]` | accurate token counting + waste findings via prompt-flamegraph |
| `pip install promptflame-optimizer[dev]` | pytest, build, twine |

With no dependencies at all, a built-in word-count estimator is used; with `prompt-flamegraph` installed, its tokenizer pipeline (tiktoken/`words`/model encodings) is used automatically.

## Quick start

```python
from prompt_optimizer import optimize

prompt = {
    "system_prompt": "You are a helpful assistant.",
    "chat_history": [...],      # OpenAI-style messages
    "tools": [...],             # tool schemas
    "rag_context": {...},       # retrieved docs
}

report = optimize(prompt, token_budget=8000, keep_last_n=20)

print(report.original_tokens, "->", report.optimized_tokens)
print(report.savings_pct, "% saved, under budget:", report.under_budget)

use_this = report.optimized   # your prompt is never mutated
```

OpenAI/Anthropic-style request bodies (`{"messages": [...], "system": ...}`) and bare message lists are auto-normalized, just like prompt-flamegraph's `normalize()` — except `tool_calls`/`tool_call_id` keys are preserved so history trimming never splits a tool-call pair.

Dry-run without applying:

```python
from prompt_optimizer import suggest

for s in suggest(prompt, token_budget=8000):
    print(f"[{s.kind}] {s.path}: {s.message} (~{s.tokens_saved} tokens)")
```

## Strategies

Applied cheap-and-lossless first, aggressive last:

| Order | Strategy | What it does | CLI flag |
|---|---|---|---|
| 1 | `whitespace` | Collapses 3+ newlines to 2, strips trailing spaces — **fenced code blocks untouched** | `--no-whitespace` |
| 2 | `dedup` | Removes exact-duplicate leaves (NFC-aware), duplicate messages, repeated list items — keeps first occurrence | `--no-dedup` |
| 3 | `minify_tools` | Drops decorative schema keys (`title`, `examples`, `$schema`, `default`), truncates descriptions to ~120 chars | `--no-minify-tools` |
| 4 | `trim_history` | Keeps the last N messages (`keep_last_n`), or trims oldest until the budget is met. Inserts a `[... N earlier messages omitted]` marker; an assistant `tool_calls` turn and its `tool_call_id` responses are an atomic group | `--no-trim-history`, `--keep-last N` |
| 5 | `trim_rag` | When over `token_budget`, drops RAG entries with the lowest keyword overlap vs the last user message | `--no-trim-rag`, `--budget` |

Safety guards: input is deep-copied (never mutated), traversal is iterative with cycle detection and a 500-level depth cap, and user strings are never parsed as JSON. `preserve=("system_prompt",)` keeps named top-level sections verbatim.

## CLI

```bash
prompt-optimizer prompt.json                     # report + optimized prompt on stdout
prompt-optimizer prompt.json -o slim.json        # optimized prompt JSON to a file
prompt-optimizer prompt.json --format json       # full report as JSON
prompt-optimizer prompt.json --budget 8000       # trim to fit (exit 3 if unmet)
prompt-optimizer prompt.json --keep-last 10      # keep last 10 history messages
prompt-optimizer prompt.json --no-dedup          # disable a strategy
prompt-optimizer prompt.json --dry-run           # suggestions only, no output
prompt-optimizer --demo                          # try it on a bloated sample
cat openai_request.json | prompt-optimizer -     # stdin works too
```

## Works great with prompt-flamegraph

```python
from prompt_flamegraph import profile_prompt
from prompt_optimizer import optimize

# see where the tokens go...
profile_prompt(prompt, output="before.html")

# ...then cut them
report = optimize(prompt, token_budget=8000)
profile_prompt(report.optimized, output="after.html")
```

When `prompt-flamegraph` is importable, `optimize()` measures with its tokenizer and surfaces `near_duplicate` findings it cannot auto-fix as informational suggestions.

## API

```python
from prompt_optimizer import (
    optimize,        # prompt dict → OptimizeReport
    suggest,         # prompt dict → list[Suggestion] (no changes)
    OptimizeConfig,  # budgets, keep_last_n, per-strategy toggles, preserve
    OptimizeReport,  # .optimized .token_savings .savings_pct .under_budget
    Suggestion,      # .kind .path .message .tokens_saved .action
)

OptimizeConfig(
    token_budget=None, keep_last_n=None,
    dedup=True, whitespace=True, minify_tools=True,
    trim_history=True, trim_rag=True,
    preserve=("system_prompt",), model=None, tokenizer=None,
)

report.to_dict()       # JSON-serializable report
report.to_markdown()   # human-readable report
```

## Source

<https://github.com/fjjjuv/prompt-optimizer>

## License

LGPL-3.0-or-later — see [LICENSE](LICENSE) and [LICENSE.GPL](LICENSE.GPL).
