Metadata-Version: 2.4
Name: pcdlint
Version: 0.3.0
Summary: Static Taint Linter for Prompt-Cache Determinism
Author: pcdlint Contributors
License-Expression: MIT
Project-URL: Homepage, https://github.com/ZhiwarSajadi/pcdlint
Project-URL: Repository, https://github.com/ZhiwarSajadi/pcdlint
Project-URL: Issues, https://github.com/ZhiwarSajadi/pcdlint/issues
Keywords: anthropic,llm,linter,openai,prompt-cache,prompt-caching,static-analysis,taint-analysis
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: rich>=13.0
Requires-Dist: tomli>=2.0; python_version < "3.11"
Provides-Extra: dev
Requires-Dist: mypy~=2.3.1; extra == "dev"
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=5.0; extra == "dev"
Requires-Dist: ruff~=0.16.9; extra == "dev"
Requires-Dist: tomli>=2.0; extra == "dev"
Dynamic: license-file

# pcdlint (Prompt-Cache Determinism Linter)

[![CI Status](https://github.com/ZhiwarSajadi/pcdlint/actions/workflows/ci.yml/badge.svg)](https://github.com/ZhiwarSajadi/pcdlint/actions/workflows/ci.yml)
[![PyPI version](https://img.shields.io/pypi/v/pcdlint)](https://pypi.org/project/pcdlint/)
[![Python versions](https://img.shields.io/pypi/pyversions/pcdlint)](https://pypi.org/project/pcdlint/)
[![License: MIT](https://img.shields.io/badge/license-MIT-green)](https://opensource.org/licenses/MIT)

A static taint analysis linter that detects code patterns silently invalidating LLM Prompt Caching (OpenAI & Anthropic SDKs).

---

## The $10,000 Bug Explained

LLM prompt caching works by storing the **prefix** of a prompt in a fast cache. When the prefix matches a cached version, the model skips re-processing the prefix — costing **up to 90% less** than a full write, depending on the provider and model (see [OpenAI's pricing](https://platform.openai.com/docs/pricing) and [Anthropic's pricing](https://platform.claude.com/docs/en/about-claude/pricing)).

The catch? The cache key is computed from the **exact byte sequence** of the prompt prefix. A single non-deterministic character at position 0 — like `datetime.now()` — changes every single byte, making the cache **completely useless**.

```
Full Write Cost:  ██████████  ($0.030 / 1K tokens)
Cache Read Cost:  █          ($0.003 / 1K tokens)  ← 90% cheaper
```

One `datetime.now()` at character 0 silently destroys cache hits, wasting money on every request.

---

## ❌ Bad Code (0% Cache Hit)

```python
from datetime import datetime

system = f"Time: {datetime.now()}\n{STATIC_RULES}"
#                                    ^^^^^^^^^^^^ PREFIX TAINTED — every request has a unique prefix

client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "system", "content": system}],  # Cache MISS every time!
)
```

## ✅ Good Code (90% Cost Reduction)

```python
from datetime import datetime

system = f"{STATIC_RULES}\nTime: {datetime.now()}"
#       ^^^^^^^^^^^^^^ Static prefix first, dynamic value at the end — cache prefix preserved

client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "system", "content": system}],  # Cache HIT — prefix bytes match!
)
```

---

## How caching differs by provider

Both providers cache a **prefix**, and neither will cache anything if the
bytes keep moving — but they decide what counts as a hit differently:

| | OpenAI | Anthropic |
|---|--------|-----------|
| Opting in | automatic | explicit `cache_control` breakpoints |
| What must match | the longest prefix that still matches | every byte up to the last breakpoint |
| Dynamic value at the *end* | partial hit on the part that still matched | — |
| Dynamic value *before* the breakpoint | partial hit | **100% miss** |

So the same code is not equally bad everywhere. OpenAI falls back to the
longest prefix that still matches, which is why `PCL001` only complains when
the dynamic value comes *before* the static text. Anthropic serves a cached
prefix only when it is byte-identical up to a `cache_control` breakpoint, so
`STATIC + f"{now}"` inside the block carrying that breakpoint throws away the
whole cache — that is what `PCL005` reports.

A `cache_control` on the *request* itself (Anthropic's automatic caching)
counts the same way: it applies the breakpoint to the last cacheable block,
so taint anywhere in `system` or `messages` is a 100% miss there too.

See [OpenAI's prompt caching guide](https://platform.openai.com/docs/guides/prompt-caching)
and [Anthropic's prompt caching docs](https://platform.claude.com/docs/en/build-with-claude/prompt-caching).

---

## Rule Reference

| Rule ID | Name | Severity | `--fix` | Description |
|---------|------|----------|---------|-------------|
| `PCL001` | `prefix-taint-injection` | ERROR | — | Dynamic value placed before static prompt text invalidates cache prefix |
| `PCL002` | `unsorted-json-in-prefix` | WARNING | ✅ | `json.dumps()` without `sort_keys=True`: key order follows how the dict was built, so two runs can disagree. A dict/list *literal* is exempt — its order is the source order. |
| `PCL003` | `set-iteration-in-prompt` | ERROR | ✅ | Python's `PYTHONHASHSEED` randomizes set iteration order across processes |
| `PCL004` | `dynamic-tools-mutation` | WARNING | — | Changing tool definition order invalidates the entire prompt cache hierarchy |
| `PCL005` | `taint-before-cache-breakpoint` | ERROR | — | Dynamic value anywhere before an Anthropic `cache_control` breakpoint, on a block or on the request itself (automatic caching) |

`PCL001`, `PCL004` and `PCL005` have no mechanical fix: only you know where
the dynamic value belongs, or what the tool order should be.

---

## Installation

```bash
pip install pcdlint
```

Or install from source with dev dependencies:

```bash
pip install -e ".[dev]"
```

## CLI Usage

```bash
# Check a single file
pcdlint check src/my_app.py

# Check multiple paths
pcdlint check src/ tests/

# Output as JSON
pcdlint check src/ --format json

# Output SARIF 2.1.0 for GitHub Code Scanning
pcdlint check src/ --format sarif > pcdlint.sarif

# Apply the mechanical fixes (PCL002 sort_keys, PCL003 sorted) in place
pcdlint check src/ --fix

# Only report findings on lines this branch changed (default ref: HEAD)
pcdlint check src/ --diff origin/main

# Fail on warnings (useful for CI)
pcdlint check src/ --fail-on-warn

# Run only some rules / skip others (overrides [tool.pcdlint])
pcdlint check src/ --select PCL001,PCL003
pcdlint check src/ --ignore PCL002

# Leave files out of the scan entirely (overrides [tool.pcdlint])
pcdlint check src/ --exclude "tests/*" --exclude "*.min.py"

# Equivalent module forms (no console script needed)
python -m pcdlint check src/
python -m pclint check src/
```

`--fix` rewrites only what it can rewrite safely, re-reads the files so the
report and the exit code describe what is on disk now, and tells you how many
findings it had to leave to you. A rewrite that would no longer parse is
reported and skipped rather than written.

`--diff` needs `git`. It compares the working tree against `REF` and keeps a
finding only when **the line the finding points at** changed — a finding on an
untouched line stays hidden even if the value it refers to was edited. Pair it
with `--fix` to repair only what a PR introduced.

A file git does not know about yet counts as wholly new, so every finding in
it is reported; files your `.gitignore` excludes are left out entirely.

### Exit codes

| Code | Meaning |
|------|---------|
| `0` | No findings |
| `1` | Findings at ERROR severity, or any WARNING when `--fail-on-warn` is set |
| `2` | A path could not be analyzed: missing, not a `.py` file, not valid UTF-8, a syntax error, or `--diff` could not run (no git, not a repository, unknown ref) |

Exit code `2` exists so a typo'd path or a broken file can never look like a clean run
in CI. Errors are printed to stderr; `--format json` output on stdout stays valid JSON.
A failed `--diff` also exits `2`: reporting "no findings" because git was
unavailable would be the linter lying about your code.

## Configuration

Rules can be turned on and off from `pyproject.toml` or the command line:

```toml
[tool.pcdlint]
select = ["PCL001", "PCL003"]   # run only these rules
ignore = ["PCL002"]             # skip these, applied after select
exclude = ["tests/*", "*.min.py"]  # never scan these at all
taint-sources = ["mypkg.jitter"]    # calls you know are nondeterministic
sinks = ["mywrapper.ask"]           # calls that take a prompt
```

The **nearest `pyproject.toml` above each analyzed file** is used, so a monorepo
can give every package its own rule set. `select` is an allowlist; `ignore` is a
denylist applied after it. `exclude` drops files a **directory walk** would
pick up, before anything is analyzed — matched against the bare file name
*and* the path relative to what you passed, so `"*.min.py"` and `"tests/*"`
both do what they look like. A file you name on the command line is always
checked.

`taint-sources` and `sinks` are the escape hatch for calls this linter
cannot know about: the first names qualified calls that are
nondeterministic (`mypkg.jitter()` taints like `uuid.uuid4()`), the second
names qualified calls that take a prompt, so everything inside one is
judged as payload. Both are taken at your word — a name is matched exactly
or on a dot boundary, and nothing else about the call has to look like an
LLM API. A typo is an error rather than a rule that silently never fires.

A folder holding a `pyvenv.cfg` is skipped, whatever it is called: `venv311/`,
`.venv-py312/` and `.tox/py311/` are caught without needing a name on the
built-in skip list. That list — packaging output and tool caches — is
`.venv`, `venv`, `node_modules`, `.git`, `__pycache__`, `build`, `dist`,
`.tox`, `.nox`, `.mypy_cache`, `.ruff_cache`, `.pytest_cache`, `.eggs`,
`htmlcov`, `site-packages` and `.ipynb_checkpoints`. `env` is deliberately
**not** on it: it is a common name for application code, and the
`pyvenv.cfg` rule already catches a virtualenv whatever it is called. A path
you name on the command line is never skipped by any of this.

`--select` and `--ignore` take a comma-separated list and are repeatable. Each
flag **replaces** the corresponding config value for that run rather than merging
with it, so `--ignore PCL002` overrides a configured `ignore` list outright.

A config with an unknown key, an unknown rule id, or broken TOML exits `2` with a
message on stderr. It never falls back to defaults — otherwise `select = ["PCL999"]`
would make the run look clean.

## Suppressing a Finding

Switch rules off on a single line with a `# pcdlint: disable` comment, placed on
the line pcdlint reports:

```python
system = f"Time: {now}\n{STATIC_RULES}"  # pcdlint: disable
```

| Form | Scope |
|------|-------|
| `... # pcdlint: disable` | that line, every rule |
| `... # pcdlint: disable=PCL001,PCL003` | that line, the named rules only |
| `# pcdlint: disable` alone, above the first statement of the file | the whole file |
| `# pcdlint: disable=PCL002` alone, above the first statement | the whole file, the named rules only |
| `# pcdlint: disable` alone, further down | the next line of code |

Only a marker written above the first statement of the file switches the
whole file off — that is where it has always lived. The same marker further
down belongs to the line of code beneath it, so putting it on its own line
above the call you mean does what you expect instead of muting every other
finding in the file. Blank and comment-only lines between the marker and its
target are skipped.

Comments are read with Python's tokenizer, so the same text inside a string
literal is data and does nothing. A marker naming a rule that does not exist
suppresses nothing — a typo surfaces the finding instead of hiding it.

The marker may sit anywhere in the comment and be written without padding, so
it can share a token with another pragma. All of these suppress:

```python
system = ...  # type: ignore  # pcdlint: disable
system = ...  # noqa: E501 pcdlint: disable
system = ...  #pcdlint:disable
```

It still has to be a complete marker: `# pcdlint: disable-all` and prose that
merely mentions the keyword do not count.

Comment the line pcdlint prints. For `PCL001` that is the line holding the
`system=...` argument, not the line where the tainted value was built.

## GitHub Actions CI Integration

```yaml
name: pcdlint
on: [push, pull_request]
jobs:
  lint:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      security-events: write   # required to upload SARIF
    steps:
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
      - uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7
        with:
          python-version: "3.12"
      - run: pip install pcdlint
      # Findings appear directly on the PR diff via Code Scanning.
      # pcdlint exits 1 for findings and 2 when it could not analyze. Only 1
      # may be absorbed here: run this step before the enforcing one, with a
      # bare `>`, and it fails on every run that has findings -- which is
      # exactly the run whose results you wanted uploaded.
      - name: Emit SARIF
        run: pcdlint check src/ --format sarif > pcdlint.sarif || test $? -eq 1
      - uses: github/codeql-action/upload-sarif@2892aa5e19bbd11bc0cff5427e3b750a04d9e3c2 # v4
        if: always()
        with:
          sarif_file: pcdlint.sarif
      - name: Enforce
        run: pcdlint check src/ --fail-on-warn
```

Code Scanning has to be enabled for the repository (free for public repos).
Add `continue-on-error: true` to the upload step if some of your repositories
do not have it enabled.

To lint only the lines a pull request changed, swap the enforcing step for:

```yaml
      - name: Enforce
        run: pcdlint check src/ --diff "origin/${{ github.base_ref }}" --fail-on-warn
```

`--diff` needs the base branch to exist locally, and a default checkout is a
shallow, single-ref clone where `origin/<base>` does not -- pcdlint would
exit 2 with "unknown revision". Give that variant a full history:

```yaml
      - uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7
        with:
          fetch-depth: 0
```

---

## How It Works

1. **AST Parsing**: Uses Python's built-in `ast` module to parse source files — no external compilers needed.
2. **Taint Tracking**: Identifies non-deterministic sources (`datetime.now()`, `uuid.uuid4()`, `os.urandom()`, etc.) and tracks their propagation through f-strings, concatenation, and `.format()`.
3. **Prefix Analysis**: Precisely determines if a tainted expression appears **before** a static prefix solid in the prompt string.
4. **Rule Matching**: Checks tainted prompts against LLM API sink calls (`chat.completions.create`, `messages.create`, etc.) and reports violations.
