Metadata-Version: 2.4
Name: cheapskate
Version: 0.1.0
Summary: Drop-in LangChain/LangGraph middleware that cuts frontier LLM API costs via tool pruning, prompt squeezing, and hybrid SLM routing.
License-Expression: MIT
License-File: LICENSE
Keywords: langchain,langgraph,llm,cost-optimization,routing,tool-pruning,prompt-compression
Author: CheapSkate Contributors
Author-email: maintainers@cheapskate.dev
Requires-Python: >=3.9,<4.0
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Typing :: Typed
Requires-Dist: langchain-core (>=0.3.0,<1.0.0)
Requires-Dist: langchain-groq (>=0.2.0,<1.0.0)
Requires-Dist: langchain-openai (>=0.2.0,<1.0.0)
Requires-Dist: loguru (>=0.7.0)
Requires-Dist: numpy (>=1.24.0)
Requires-Dist: tiktoken (>=0.7.0)
Project-URL: Bug Tracker, https://github.com/cheapskate-ai/cheapskate/issues
Project-URL: Documentation, https://github.com/cheapskate-ai/cheapskate#readme
Project-URL: Homepage, https://github.com/cheapskate-ai/cheapskate
Project-URL: Repository, https://github.com/cheapskate-ai/cheapskate
Description-Content-Type: text/markdown

# cheapskate

**Drop-in cost control for LangChain and LangGraph agents.**

`cheapskate` sits in front of expensive frontier models (GPT-4o, Claude-class APIs) and automatically applies three levers that cut token waste and route easy work to cheap SLMs — without rewriting your agent graph.

Built for teams that already ship tool-calling agents and want lower bills without giving up quality on hard tasks.

---

## Why it exists

Agent stacks get expensive for boring reasons:

- every turn re-sends bloated tool schemas
- prompts accumulate filler and repeated boilerplate
- simple classify / summarize / translate calls still hit GPT-4o prices

`cheapskate` intercepts those calls as a LangChain `BaseChatModel`, so it works as a middleware layer inside existing chains and stateful LangGraph workflows.

---

## Core features

### 1. Advanced tool pruning
Keeps only the tool schemas that matter for the current turn.

- Groups tools into **namespaces** (weather, finance, ops, …)
- Scores relevance with **context-aware vector similarity** over full message history, including prior tool results / scratchpad text
- Uses numpy cosine similarity — not naive keyword matching alone
- Optional **namespace classifier hook** (plug in an SLM or custom ranker)
- Hard guarantee: tools listed in `always_keep` are never pruned

**Effect:** fewer tools in the prompt → fewer input tokens on every frontier call.

### 2. Prompt squeezing
Compresses conversational fluff while protecting structured payloads.

- Operates on a **shallow copy** of messages so LangGraph checkpointers keep an uncorrupted history
- Scrubs high-frequency boilerplate and filler phrases
- Uses **token-density** heuristics (via `tiktoken`) to truncate low-signal older turns
- Leaves valid JSON blocks inside message text intact

**Effect:** smaller context windows without silently breaking tool arguments or stored state.

### 3. Hybrid SLM routing
Sends easy, deterministic work to a cheap secondary model (Groq / Together-style Llama-class SLMs).

- Lightweight heuristics decide when a task is “simple enough”
- Hard / ambiguous / tool-heavy work stays on the primary frontier model
- **Failsafe path:** if the SLM errors, times out, or returns invalid structured output, cheapskate logs the failure and immediately retries on the primary model with the **original unmodified payload**

**Effect:** most of the spend reduction on repetitive turns, without single-point-of-failure routing.

### 4. Observability built in
`CheapSkateCallbackHandler` + router metrics track:

- baseline vs optimized token counts
- prune / compression / routing decisions
- failover events

---

## Quick start

```bash
pip install cheapskate
```

```python
from langchain_openai import ChatOpenAI
from langchain_groq import ChatGroq
from cheapskate import CheapSkateRouter, ToolNamespace

primary = ChatOpenAI(model="gpt-4o")
secondary = ChatGroq(model="llama-3.2-3b-preview")

router = CheapSkateRouter(
    primary_model=primary,
    secondary_model=secondary,
    always_keep_tools=["search_docs", "calculator"],
    tool_namespaces=[
        ToolNamespace(
            name="weather",
            tools=("get_weather", "get_forecast"),
            description="weather forecast temperature",
        ),
        ToolNamespace(
            name="finance",
            tools=("get_stock_price", "list_portfolios"),
            description="stocks portfolios equity",
        ),
    ],
)

response = router.invoke(
    [{"role": "user", "content": "Summarize this in one sentence: rain is likely Saturday."}]
)
print(response.content)
print(router.metrics.summary)
```

Drop `CheapSkateRouter` anywhere you currently pass a chat model in LangChain / LangGraph.

---

## Architecture

```text
Agent / LangGraph node
        │
        ▼
┌───────────────────────────┐
│     CheapSkateRouter      │
│  1. prune tool schemas    │
│  2. squeeze prompt tokens │
│  3. route simple → SLM    │
│  4. failover → primary    │
└─────────────┬─────────────┘
              │
     ┌────────┴────────┐
     ▼                 ▼
 cheap SLM        frontier LLM
 (Groq/Together)  (GPT-4o / Claude)
```

---

## Benchmarking without API keys

You can measure pruning, compression, routing, and **synthetic cost impact** offline:

```bash
poetry install --with dev
poetry run python benchmarks/offline_bench.py
```

This uses mock chat models + `tiktoken` + published list-price assumptions. No OpenAI / Groq keys required. Results are written to `benchmarks/offline_results.json`.

Live quality-vs-cost A/B against production traffic still needs real keys — the offline bench validates the middleware mechanics and estimates spend deltas.

---

## What is production-ready today

| Capability | Status |
|---|---|
| tiktoken token accounting | Complete |
| Tool pruning + `always_keep` | Complete |
| Prompt compression on shallow copies | Complete |
| Hybrid routing heuristics | Complete |
| SLM → primary failsafe | Complete |
| Sync + async `_generate` / `_agenerate` | Complete |
| Metrics / callback handler | Complete |
| Default hosted SLM namespace classifier | Hook only (bring your own) |
| Neural embedding provider | Local hash embeddings by default |

The common “60–80% savings” range is a **target envelope** for tool-heavy agents with lots of prompt fluff and easy turns. Real savings depend on tool cardinality, history length, and how often work is SLM-eligible. Run the offline bench, then validate on your traffic.

---

## Development

```bash
poetry install --with dev
poetry run pytest -q
./publish_pipeline.sh   # needs Artifactory/PyPI credentials to publish
```

---

## License

MIT

