Metadata-Version: 2.4
Name: mountain-of-helicon
Version: 0.2.0
Summary: Check AGENTS.md and CLAUDE.md claims against the repository on disk
Author: Oscar Morkeeth
License-Expression: MIT
Project-URL: Homepage, https://github.com/Morkeeth/mountain-of-helicon
Project-URL: Repository, https://github.com/Morkeeth/mountain-of-helicon
Project-URL: Issues, https://github.com/Morkeeth/mountain-of-helicon/issues
Keywords: ai-agents,agent-memory,claude-code,mcp,governance
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: openai>=1.30.0
Requires-Dist: fastapi>=0.111.0
Requires-Dist: uvicorn>=0.30.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: gitpython>=3.1.0
Requires-Dist: numpy>=1.26
Provides-Extra: embeddings
Requires-Dist: sentence-transformers>=3.0; extra == "embeddings"
Provides-Extra: test
Requires-Dist: pytest>=8.0; extra == "test"
Requires-Dist: httpx<1.0,>=0.27; extra == "test"
Dynamic: license-file

# Mountain of Helicon

**Your `AGENTS.md` is lying to your coding agent.** It points at files that moved,
commands that no longer exist, paths the repo reorganized away. Your agent loads
those rules as fact at the start of every session. Nothing tells you until
something breaks.

```bash
pip install mountain-of-helicon
helicon review .
```

What it printed on our own public repo, `MorkeethHQ/world-relay`, from a fresh clone of
`main` on 2026-09-03:

```
  ✗ 1 graded claim contradicted by repo evidence.

    ✗ AGENTS.md:19  points at src/__tests__/e2e-api.test.ts  — not in this repo

  GRADE B   ·   35 claim checks, 1 contradiction
  1 instruction claim conflicts with observed repo evidence.
```

The test moved to `src/lib/__tests__/e2e-api.test.ts`; the rule that tells the agent how to run it did not
move with it. The same command on `anthropics/anthropic-cookbook`, same day, prints
**GRADE A, 6 references checked, 0 broken**. An earlier build of this tool graded
that repo D on four pointers that all resolved; the fix and its fixtures are in
[`docs/HELICON-ON-OUR-OWN-BOARD-2026-09-03.md`](docs/HELICON-ON-OUR-OWN-BOARD-2026-09-03.md).

No API key. No config file. No upload. It reads the agent rules your repo already
commits (`AGENTS.md`, `CLAUDE.md`, `.cursorrules`) and checks every pointer against
the tree on disk. External host paths and create-on-demand outputs are shown as
unverified and excluded from the grade; they cannot pad or lower it. Exit code is
non-zero when a graded repo-local claim fails, so it drops straight into CI.

**What the install pulls.** The review needs no API key, no database and no config
when it runs. The install is not that small. `pip install mountain-of-helicon` also
installs openai, fastapi, uvicorn, numpy, gitpython and pyyaml, because other
`helicon` commands (the local web app, the memory store) use them. `helicon review`
does not load any of them.

Second command: `helicon witness` grades your agent's last session instead of your
repo: which of its "done / passes / fixed" claims have tool evidence in the same
trace. On the author's own last 20 sessions the median verified-claims share was
**0.42** (measured 2026-09-02; 10 of the 20 carried a checkable claim).

## The other two doors: your agent's memory, and its claims

The review above reads a repo. The same engine reads a memory directory and a session
transcript, with no key and no config:

```bash
helicon truth ~/.claude --recursive   # which of your agent's documents are lying, and why
```

**No API key. No database. No config at run time.** It reads a directory and returns a ranked report,
and every row cites the line it fired on. On this machine, that first run reads:

```
1182 files scanned · 628 carry a staleness/rot signal · 554 clean
  1   39   GOLDEN_RULES.md
        +22  expired dated claim (108d past, 2026-05-15)  -> "Close-Out - May 15"
```

```bash
helicon witness           # your last agent session: every claim vs its evidence
helicon setup --audit     # local Claude, Cursor and Codex setup evidence
```

For a specific project, run `helicon setup --audit --project /path/to/repo`.

The Setup screen also includes an independent, read-only ZUP project-state review.
It compares local event receipts, canonical board identities, and queue actions.
Duplicate identities, missing receipts, inconsistent state, and actions on
settled phases are findings. Missing stores are unmeasured. A clean comparison
does not prove every project is current or that memory improves outcomes.
The source is `ZUP_HOME` (default `~/.zen`); no private records are published.

The Setup page also includes an index-and-memory operating review: live sources,
scan errors, live embedding coverage and model mix, exact duplicate hashes,
recorded retrieval use, and live keyword-search smoke probes. Expand each check
for its rows, query and next action. Relevance, contradiction quality and causal
benefit are explicitly unmeasured here; a working search is not a correct answer.
For recorded results, `helicon outcomes` shows accepted, rework, rollback and
missing acceptance decisions, with the recorded-run denominator. It does not
grade all agent work. `--json` exports the `helicon.outcomes/1` contract for ZUP.
Use `helicon outcomes --save-baseline /local/path/baseline.json` to freeze a
reading in a new file; later use `--baseline /local/path/baseline.json` to compare
the same run IDs. New runs are excluded; missing old runs mark the comparison
incomplete. Acceptance changes are not evidence of causal setup benefit.

Add `--json` for the `helicon.setup-audit/1` contract consumed by ZUP.
The audit reads instruction and skill files, reports their hashes, and checks
inspectable Claude hook routes and post-compaction vault references. It neither
executes hooks nor changes configuration. Reports contain local paths and hashes,
not instruction bodies or configuration secrets; review paths before sharing.

Discovered files are candidates, not proof they reached a model. Effective context,
plugin activation, cloud settings, and skill benefit remain unmeasured. The older
`helicon setup` census remains available, but its limited file count no longer
passes as a measurement of total loaded context.

One real catch, from a real transcript, in under a minute:

```
[NO-EVIDENCE ] L1127: "I ran the full pytest suite and all tests pass."
               → no tool call in this run could support it
```

Your agent said it. The trace doesn't back it. Now you know before you merge.

**Your CLAUDE.md is lying to your agent, and nothing tells you.**

It says "see `docs/architecture.md`" after that file moved. It routes to a
directory the monorepo split apart. It warns about a rail that was retired six
weeks ago. Your agent loads all of it as fact at the start of every session, acts
on it, and neither of you finds out.

Mountain of Helicon runs the claims in those files against the repository in
front of it, before the work starts. Every contradiction it reports carries the
command and the stdout that proved it — so you are ruling on evidence, not on a
model's opinion about your docs.

### One line, no clone

Review any repo's agent setup in one command — no clone, no key, no LLM:

```bash
uvx --from mountain-of-helicon helicon-review ~/your-repo
# or:  pipx run --spec mountain-of-helicon helicon-review ~/your-repo
```

It reads that repo's `CLAUDE.md` / `AGENTS.md` / `.cursorrules`, checks every
pointer, command, and version claim against the actual tree, and prints a graded
verdict with the exact `file:line` of each broken reference. Exit code is non-zero
when the setup lies to its agent, so it drops straight into CI.

> Runs the existence + version tier by default. The execute-and-compare wedge — it
> *runs* a documented test/build command and grades the doc's "this passes" claim
> against the real exit code — is opt-in (`HELICON_EXECUTE=1`), because running a
> stranger's code should be a choice.

### From a clone (the full lab)

```bash
git clone https://github.com/Morkeeth/mountain-of-helicon.git
cd mountain-of-helicon
python3 scripts/check_python.py     # not optional; see below
python3 -m pip install -e .

helicon review ~/your-repo          # the front door: graded, evidence-backed review
helicon ci --path ~/your-repo       # your repo. no key, no config, no init.
```

That last line is the product. It reads the `CLAUDE.md` / `AGENTS.md` /
`.cursorrules` / `.clinerules` / copilot-instructions **your** repository commits,
probes each claim against **your** code, and prints every contradiction with the
command and the stdout that proved it. Nothing is uploaded, no key is asked for, and
no configuration file is written.

It is fast because there is nothing heavy in the path — no model download, no build
step, no index to warm. The scan is `git` and subprocess calls, so on a cold clone it
finishes in seconds; two different machines measured 3.5s and 2s for the scan itself.
Time it on yours rather than trusting either number.

If you would rather watch it work on a known-convicted repository before pointing it
at your own:

```bash
bash scripts/demo.sh                # a real gate firing, with evidence, no key
```

The demo scans named public repositories, shows contradictions with the probe and
stdout that proved each one, installs the preflight hook into a throwaway settings
file, and records an explicit override. It touches neither your real Claude
settings nor any memory store.

Local-first. No API key for the deterministic checks. It **warns by default** and
never blocks unless you opt in, because a preflight that wedges a terminal gets
uninstalled. It can also audit longer-lived memory, let a human rule on
contradictions, and compile those rulings into policy your agents query over CLI
or MCP.

**What it does not do:** it settles what the filesystem can settle — a named path
that is gone, a quoted command's output, a retired capability. It cannot tell you
whether a sentence is *true*. Silence is not a clean bill of health; it means no
executable probe could bind.

## The measured finding

We ran the frozen 591-repository corpus and scored 577 current default branches;
14 exclusions are named. After removing projection duplicates, enforcing the
existing-file evidence invariant, and hand-verifying all 30 mechanical survivors,
**6 repositories (1.04%) contained a sendable doc-vs-code contradiction**.
Finding-level precision was 9/30. The rejected rows and reasons are part of the
result, not a footnote.

The frozen corpus, commands, stdout, exclusions, and hand-verification ledger are
in [`docs/agent-context-report-2026-08.md`](docs/agent-context-report-2026-08.md).
Release gates and intentionally deferred work are tracked in
[`LAUNCH_ROADMAP.md`](LAUNCH_ROADMAP.md).

Build notes, night-run logs and review prompts live in
[`docs/archive/`](docs/archive/) — kept, not deleted, just out of the front door.

## Installing on a Mac

**It took six cold clones before one ran clean.** Not six errors found by reading
the code — six real clones from GitHub into `/tmp`, installed into an empty
`HOME`, driven with the stock macOS interpreter, following the block above
exactly. Attempts 1, 3 and 5 each surfaced something the previous one had not:
output that greeted a stranger by the author's name, `pip install` blaming the
repo for a missing `setup.py` that was really a three-year-old pip, and a server
that failed to bind and still printed `open http://127.0.0.1:8420`. The suite was
green through all of it, because none of those are things a test suite is looking
at. If you hit a seventh, that is a bug and worth an issue.

**Run `check_python.py` first; it is not a formality.** Python 3.10+ is required,
and a stock Mac fails this two different ways, neither of which names the real
cause. `/usr/bin/python3` is 3.9 and dies on a PEP 604 annotation
(`TypeError: unsupported operand type(s) for |`). Its bundled pip is 21.2.4,
which predates PEP 660 and refuses with
`ERROR: File "setup.py" or "setup.cfg" not found` — that one reads as *this repo
is missing a file* rather than *your pip is three years old*. The preflight
checks both and prints the exact command for each.

**Bring your own Qwen key (BYOK).** Get one free on the [Alibaba Cloud Model Studio free tier](https://www.alibabacloud.com/en/product/modelstudio), set `QWEN_API_KEY` or put it in `~/.helicon/config.json`. `helicon init` keeps configuration and the SQLite store under `~/.helicon/`, never inside the installed package. **Keyless degrade:** without a key every deterministic test still runs; only the two LLM-judged tests (Contradiction, Grounding) switch off -- the battery says so instead of faking a verdict.

## What ships

- **Doorway:** `helicon sweep` checks agent-rules files against repository
  reality; `helicon doorway install` adds a reversible Claude Code preflight.
- **Memory governance:** a 13-class rot exam, human rulings, receipts, undo, and
  Golden Rules.
- **Agent access:** local MCP exposes 25 tools, plus an authenticated remote endpoint with a narrower allowlist.
- **Connectors:** Claude Code, Cursor and Cursor Cloud exports, git, Obsidian,
  agent rules, ChatGPT exports, Mem0, Letta, Graphiti, and LifeOS adapters.
- **Dashboard:** Doorway, Rulings, governed runs, memory health, and the deeper
  Lab surfaces.

## Visual demo

The web bundle is generated, never committed stale. From a source checkout:

```bash
helicon demo
```

On first run this installs/builds the dashboard with npm, seeds a labelled
19-memory demo under `~/.helicon/demo`, and serves it only on
`http://127.0.0.1:8420/#findings`. No personal connector runs and no API key is
required.

## Use it on your own repo

One command. No key, no config, no `init`. Run it inside a repository you own:

```bash
helicon ci
```

Use `--path <dir>` to check a repository you are not standing in. It reads that
repo's committed `CLAUDE.md` / `AGENTS.md` / `.cursorrules` / `.clinerules` /
copilot-instructions, probes each claim against the repository in front of it, and
prints every contradiction with the command and the stdout that proved it.

```
  CONTRADICTED  CLAUDE.md:9   [path]
     claim   "Read `MASTER_GUIDE.md` for the operating model."
     probe   $ git ls-files -- MASTER_GUIDE.md
     output  (no output)
     why     the doc names MASTER_GUIDE.md; git tracks no such file and it is not on disk
```

If it prints nothing, no executable probe could bind. That is not a clean bill of
health, and the section above says why.

### Then, if you want the rest

The store, the dashboard, the rulings and the memory audit need a configuration
file, so they start with `helicon init`. Running `helicon scan` before `init` will
tell you the config is missing rather than do anything.

```bash
helicon init                       # writes ~/.helicon/config.json
helicon scan                       # extract memory from your configured sources
helicon doctor                     # PATH, config, key, DB, last scan
helicon audit
helicon check "what am I working on"
helicon serve
```

Qwen is optional and BYOK. Without a key, deterministic checks continue and the
two LLM-judged checks report themselves unavailable rather than fabricating a
verdict. Semantic embeddings are an optional install; the core remains slim.

## CI for agent memory (GitHub Action)

The same `helicon ci` from the section above also runs in CI, so a pull request that drifts your agent's instruction files fails the build — CI for memory, literally. It scans a repo's committed `CLAUDE.md` / `AGENTS.md` / `.cursorrules` / `.clinerules` / copilot-instructions, runs 13 documented failure classes through the 13-class deterministic exam (no key, no torch, no LLM), emits GitHub annotations + a job-summary table, and exits non-zero on rot. R13 goes further than reading: it runs a probe against the repo's own running code and reports which sentences the system contradicts.

```yaml
# .github/workflows/memory-ci.yml
name: memory-ci
on: [push, pull_request]
jobs:
  rot-exam:
    runs-on: ubuntu-latest
    steps:
      - uses: Morkeeth/mountain-of-helicon@main
        with:
          fail-on: rot   # or 'none' for report-only
```

Locally it's the same one command: `helicon ci`. This repo dogfoods the exam in
report-only mode (`--fail-on none`) so known R6 findings remain visible without
making unrelated pull requests permanently red. Teams that have ruled their
baseline clean should use the Action's default `fail-on: rot`.

## The live doorway (a warning backed by executable proof)

Everything above produces a verdict. The doorway puts it where work begins: a
Claude Code `UserPromptSubmit` hook warns when the repository disproves loaded
instructions, naming the offending lines and executable evidence in the terminal.
Warning is the default because a preflight that wedges a terminal gets removed.
Teams that explicitly want enforcement can set `HELICON_GATE_MODE=block`.

```bash
helicon board                 # every repo under ~/CODE and what it loads into an agent
helicon board --repo <name>   # every loaded line, with its probe verdict
helicon doorway install       # wire the gate into ~/.claude/settings.json (backup + diff + confirm)
helicon doorway install --uninstall   # remove exactly what it added
helicon hook --print-config   # the settings.json snippet (never auto-installed)
helicon receipt <session>     # did the harness actually RECEIVE the injection?
```

The gate a stranger installs is **keyless and config-free**: `helicon doorway install`
writes one `UserPromptSubmit` hook (shown as a diff, backed up first, idempotent,
and exactly reversible), and the hook — `python3 -m helicon doorway gate` — needs no
`config.json` to run. On the next prompt in any repo whose loaded docs its own code
disproves, the warning appears in your terminal and the run continues. Warnings and
explicit overrides log into your configured store (so they show up in `helicon runs`
/ the dashboard), and fall back to a standalone `~/.helicon` store for a stranger
who has no config.

Three rules it obeys, all from the same law:

- **Machine-evidenced.** A `CONTRADICTED` verdict came from a probe that executed and
  disagreed. The operator can correct/demote the line, continue after the warning,
  or explicitly record an override reason.
- **Cold lines never block.** Demoting a line keeps it forever and loads nothing, so
  it cannot poison a run and must not stop one. `--demote` is a real exit, not advice.
- **Fail open, loudly.** Any error lets the prompt through; a doorway that bricks a
  terminal gets uninstalled, and then it governs nothing. Logged warning/override
  events make intervention inspectable — the absence of one proves nothing.

`helicon receipt` is the honest half of delivery. Every other step can be satisfied by
a row Helicon itself wrote; this one opens the transcript the **harness** wrote and
looks for a content-derived token, ruling `RECEIVED` / `NOT_FOUND` / `UNVERIFIABLE`.
UNVERIFIABLE is a verdict, never rounded up. Injections are checked against the
~32k context-rot onset first and trimmed if they would exceed it — a memory tool that
quietly causes the rot it detects is the joke writing itself.

## Headline Features

- **`helicon snapshot`** -- regression tests for retrieved context. Capture what a task retrieves today; `snapshot check` fails when tomorrow's retrieval drifts. CI for memory.
- **`helicon check "<task>"`** -- context-quality battery on what a task retrieves: Relevance, Freshness, Redundancy, Thinness, Expiry (deterministic) + Contradiction, Grounding (judged live by Qwen). Verdict: HEALTHY / DEGRADED / BROKEN. Every verdict prints the age of the last scan, because a DEGRADED verdict is uninterpretable if the scan itself is stale. `--json` for scripts and CI.
- **`helicon reconcile`** -- timely forgetting. Re-scans sources and retires memories reality no longer contains (dry-run by default, never touches human decisions). On the live DB it retired 20 superseded memories in its first run.
- **`helicon fix-skills`** -- write-back: Qwen writes missing descriptions into your agent skill files (dry-run by default, `.bak` backups). It fixed 7 of this project's own skills.
- **`helicon doctor`** -- five checks (PATH, config, key, DB, last scan), exit 1 on failure. The front door to a daily loop.
- **`helicon rule "<natural language>"`** -- prompted rules. Qwen compiles your sentence to a restricted predicate (whitelisted fields, never code); before approval you see coverage, samples, empirical precision against YOUR past decisions, and conflicts with other rules. One approved rule governs hundreds of items; applied rules are never counted as human evidence.
- **The regret ledger** -- killed memories become a ghost list (LeCaR cache-eviction mechanics). When retrieval wants one back, a time-decayed regret event blames the exact decision that killed it, and FINDINGS shows "you retired this, retrieval wanted it 2x since -- restore?". Wrong forgetting is measured, not assumed.
- **`helicon_flag` over MCP** -- point-of-use correction. Injected memories carry id + last_verified + used_count; the agent (or you, through it) flags stale/wrong/useful in one call. Flags become findings the human confirms -- the agent proposes, it never deletes.

## Three Layers

**Layer 1 -- Extraction.** Pluggable connectors cover Claude Code, Cursor and Cursor Cloud exports, Obsidian, git history, ChatGPT exports, agent rules, LifeOS, Letta MemFS, Graphiti, and Mem0. Rewritten and expiring Mem0 memories carry their temporal fields into freshness tests. Agent *rules* files (`CLAUDE.md`, `AGENTS.md`, `.cursorrules`) are split into section-level memories so regression catches one section drifting. Every item becomes a **HeliconCube**: a versioned memory unit with source, confidence, content hash, review status, and decay parameters. A novelty gate prevents redundant storage.

**Layer 2 -- Review pattern learning.** Weibull forgetting curves with per-type shape (cliff decay for code, long tail for decisions). Auto-triage derives kill/approve rules from HUMAN reviews only -- its own decisions are excluded so it cannot reinforce its own echo. On its first run it handled 585 of the 1,268 memories the store held at that time autonomously. Spin detection, kill prediction, Helicon Score.

**Layer 3 -- Meta-audit.** The system audits its own stored patterns: temporal staleness ("this week" in a 27-day-old file), factual contradictions (Qwen-judged), decay, pattern staleness, anti-confabulation challenges. The human reviews the memory review.

## Qwen Cloud API usage (where the LLM is load-bearing)

| Tier | Model | Used for |
|------|-------|----------|
| fast | `qwen3.6-flash` | Memory summarization, novelty gate, skill descriptions |
| default | `qwen3.6-plus` | Battery judging (Contradiction, Grounding), factual audit, Next Moves |
| deep | `qwen3.7-max` | Consolidation synthesis, optimization reports |
| retrieval | `text-embedding-v4` | Dense vectors (1024-dim) for hybrid + semantic search |
| retrieval | `qwen3-rerank` | Two-stage rerank over RRF-fused candidates |

All calls go through the OpenAI-compatible endpoint `https://dashscope-intl.aliyuncs.com/compatible-mode/v1` with a per-call SQLite response cache and per-operation cost tracking (`/api/tokens`). The two subjective battery tests are judged live and tagged `(qwen)` in output; if the judge call fails, the battery falls back to deterministic-only rather than fabricating a verdict.

## MCP Server (25 tools)

Agents audit their own memory mid-conversation. Add to `.claude.json`:

```json
{
  "mcpServers": {
    "helicon": { "command": "helicon", "args": ["mcp"], "cwd": "/path/to/mountain-of-helicon" }
  }
}
```

| Tool | Description |
|------|-------------|
| `helicon_health` | Memory score and stats |
| `helicon_stale` | Decayed memories below threshold |
| `helicon_search` | Hybrid FTS5 + semantic search |
| `helicon_contradictions` | Active factual conflicts |
| `helicon_recent_reviews` | What the human approved/killed |
| `helicon_patterns` | Learned behavioral patterns |
| `helicon_guard` | Check a proposed claim against the compiled law *before* writing it: `blocked` / `warn` / `clean` |
| `helicon_ask` | Guarded retrieve — what is safe to believe about a topic: the ruled-true answer + retrieved context split into safe vs. ruled-wrong |
| `helicon_brief` | The morning brief — all five pillars in one call: truth, continuity, direction, reflection, calm |
| `helicon_portrait` | Grounded portrait of what the record shows about the person, plus its health |
| `helicon_context` | Proactive memory injection for a task -- every memory carries its id, last_verified, used_count |
| `helicon_flag` | Point-of-use correction: flag a memory stale/wrong/useful by id; stale/wrong become findings the human confirms |
| `helicon_playbook` | Task playbooks from review patterns |
| `helicon_compile` | Compile reviewed memory to injectable files |
| `helicon_triage` | Trigger auto-triage |
| `helicon_prompt_gate` | Gate an execution prompt through a Wager -- approves only after a human accepted a BUILD or REPAIR move, else abstains |
| `helicon_capture_launch` | Freeze the acceptance test and context packet before implementation starts |
| `helicon_capture_closeout` | Close a run with real artifacts and a real verification receipt |
| `helicon_workgraph_trace` | Join one work card to its task run, context, memory, skills and evidence |
| `helicon_workgraph_attention` | Name the missing graph edge -- link_run, freeze_context, attach_artifact, choose_move |
| `helicon_workgraph_learning` | Withhold recommendations until real resolved outcomes accumulate |
| `helicon_workgraph_review_skill` | Record the skill version actually loaded, hashed over its bytes |
| `helicon_consolidate` | Run a consolidation (sleep) cycle |
| `helicon_context_packet_inspect` | Inspect a named run's local source-packet receipt; does not deliver content |
| `helicon_context_packet_consume` | Return the exact reviewed project sources to the named local run and record receipt, not compliance |

The full JSON-RPC 2.0 handshake (initialize, tools/list, tools/call) is exercised in the receipts; `helicon mcp` runs the server on stdio, so the bare CLI never silently becomes a server.

### Remote MCP for cloud agents

`helicon serve` also exposes a stateless MCP endpoint at `/mcp` when a dedicated
token is configured:

```bash
export HELICON_MCP_TOKEN="$(python3 -c 'import secrets; print(secrets.token_urlsafe(32))')"
HELICON_CONFIG=/path/to/config.json helicon serve
```

Configure the remote client with `https://your-helicon-host/mcp` and send
`Authorization: Bearer <token>`. The endpoint refuses to start with a token
shorter than 32 characters, rejects requests over 1 MiB, serializes access to
the SQLite connection, and returns `Cache-Control: no-store`.

Remote access deliberately exposes the agent workflow, including context,
guard, ask, and point-of-use flags, but not `helicon_compile`,
`helicon_triage`, or `helicon_consolidate`. Those maintenance tools can write
host files or mutate the store in bulk and remain available only through the
local stdio transport.

The built-in server does not terminate TLS. Put it behind HTTPS on a private
network or an authenticated reverse proxy; never send the bearer token over
plain HTTP and never expose a personal memory store directly on a public IP.
`HELICON_PASSWORD` protects dashboard API routes and is intentionally separate
from `HELICON_MCP_TOKEN`.

### Feed Cursor Cloud runs back into memory

The `cursor-cloud` connector reads local export bundles containing
`index.json` and per-run `transcript.json`, `diff-metadata.json`, and
`events.json` files:

```json
{
  "connectors": {
    "cursor-cloud": {
      "enabled": true,
      "export_dir": "~/Downloads/cursor-cloud-agent-transcripts",
      "include_text": false
    }
  }
}
```

It selects the newest export for each stable cloud-agent id and produces one
idempotent session summary: repository, branch, model, status, message/tool
counts, observed tool failures, events, and diff/PR outcome. Metadata-only is
the default because raw exports may contain private prompts, reasoning,
terminal commands, file contents, diffs, and credentials.

Set `include_text` to `true` only when conversation prose belongs in the memory
store. That opt-in includes bounded user and final-assistant text with common
token patterns redacted. Reasoning, tool arguments, terminal output, file
contents, search results, and diffs are never ingested.

## CLI (77 commands)

`init` `scan` `reconcile` `fix-skills` `serve` `demo` `triage` `review` `route` `score-runs` `runs` `run` `hook` `receipt` `judge-bench` `bench` `attribute` `move` `leaderboard` `snapshot` `lens` `taste` `check` `report` `read` `audit` `consistency` `registry` `checkouts` `volatility` `unreviewed` `fleet` `queue` `guard` `ask` `teach` `brief` `board` `sweep` `doorway` `repair` `ci` `policy` `evolve` `wager` `capture` `lift` `resolve` `watch` `alias` `rule` `doctor` `export` `mcp` `score` `stack` `setup` `outcomes` `witness` `skills-review` `optimize` `eval` `embed` `playbooks` `reflect` `compile` `consolidate` `eval-consolidation` `complaints` `overboard` `ledger` `measure` `magnet` `measurement-bench` `review-queue` `science` `truth`

Four of them answer to a second name, kept working so older muscle memory doesn't break: `battery` = `check`, `rot` = `audit`, `heal` = `repair`, `gold` = `policy`. Aliases, not extra commands, so they are not counted above.

`helicon truth` is the **stranger-facing cold path**: point it at any agent memory directory (Claude Code, Cursor, Cline, Obsidian vault) and get a ranked, evidence-cited staleness+rot report — no Helicon DB, no API key, no LLM. Deterministic; every row cites the line it fired on.

`helicon route` turns output-verification into a **model-routing recommendation**: it reads the eval store — the verified verdicts `review-queue --terminals` produced — and ranks models by Wilson-scored verified-pass-rate per task-class, with sample size and confidence attached. The model is attributed from the git co-author trailer of the commits that produced the output; the outcome is a real reality-check, never a guess. Below a sample threshold it says *insufficient evidence*, never a fabricated number. `helicon route --record --run` builds the evidence first. See [docs/ROUTE.md](docs/ROUTE.md).

`helicon score-runs` and `helicon runs` score whole RUNS, the same output-verification one level up and made cost-aware: `score = verified yield / cost - damage`. Cost comes from the real transcript token usage, yield from the `review-queue --terminals` verdicts, damage from an incident flag. Every term traces to a real source; nothing is vibed. `score-runs --card` cuts one run card, `runs` renders the scored history, `runs --suggest` reads what to run next off it. See [docs/RUNS.md](docs/RUNS.md).

`helicon judge-bench` benchmarks Qwen as the memory-rot judge against the operator's own human rulings, and (with an OpenRouter key) against GPT/Claude on the same probes. Real result (run #2, 26 probes, 13 real contradictions + 13 controls, ground truth = the operator's own rulings): **qwen3.6-plus ties anthropic/claude-sonnet-5 at 0.962 accuracy and beats openai/gpt-5 (0.808)**, at $0.00444 per run and in 54s against Sonnet's 144s and GPT's 245s. qwen3.6-flash holds 0.923 for $0.00167, 8x cheaper than qwen3.7-max for the same specificity. And **every model, Qwen and Claude and GPT alike, missed `unit-drift`**: some rot is a domain ruling rather than a logical contradiction, and no judge at any price catches it. That is what the human-ruling layer exists for. Reproduce: `helicon judge-bench --set all --save` (needs `openrouter_api_key` for the competitors); the Judge tab reads the saved run, and renders an unrun bench as unrun. See [docs/ROUTE.md](docs/ROUTE.md).

`helicon move` is the context-mover: read memory from one platform, VERIFY each item (freshness, and with `--verify-contradictions` the Qwen judge), and render the survivors into another platform's native format (`claude-code` / `cursor` / `markdown`). Memory moves verified, never blindly; held-back items are listed with why. Dry-run by default; `--apply` backs up the target first.

`helicon leaderboard` is the population-scale version: it reads git history across many repos (where multiple models and harnesses actually co-authored commits) and ranks models by how often their commits SURVIVE vs get REVERTED, Wilson-scored. Execution-free (git only, so it is bounded and cannot freeze a machine); the revert is the honest failure signal. On 927 attributed commits across 25 local repos it already separates opus-4.6 / opus-4.8 / fable-5 / cursor by reliability.

`helicon repair` runs **the self-healing audit loop** — the thing no retriever can do. It scores the four truth gates (freshness / volatility / consistency / retrieval) on a store, surfaces each drift with its cross-source evidence, proposes a repair (retire the stale memory, move a fast fact to the live layer) as a diff you accept, applies the accepted ones, and re-scores so the gates visibly move. `helicon repair --demo` runs it on a seeded, universally-legible store (the classic "I told my agent I'm vegetarian, then started eating chicken again — it never updated" contradiction, plus a stale goal and a fast fact); `--apply` closes the loop.

`helicon audit` runs **the rot exam**: the 13 documented memory-failure classes in [ROT.md](ROT.md) checked live against your real store -- deterministic, zero LLM calls, free to run daily. On this repo's own store it currently finds rot in several of 13 classes and says so — and as of Jul 5 all classes are fully tested, 0 partial.

`helicon watch` makes the exam ambient: scan + selectors + rot exam on a timer (`helicon watch --install` writes the crontab line, every 6h), diffed against the last run. You get a macOS notification and a `drift-report.md` only when something NEW rots — no news, no noise. First run baselines silently.

`helicon policy` compiles **GOLDEN RULES**: the stack's law, built from your rulings, dismissal precedents, approved triage rules, declared renames, canonical sources and standing feedback — every rule with its provenance (a rule without provenance is a vibe). `--inject` writes it to `~/.claude/GOLDEN_RULES.md` (dry-run default, `.bak` kept) so every session can obey it. `helicon evolve` is the night command: scan, every selector, the exam, a gold recompile, and the morning delta — what your stack learned while you slept.

`helicon report` prints a **MemoryAgent Compliance Report**: the track's four sub-goals (efficient storage/retrieval, timely forgetting, recall under limited context windows, cross-session accuracy) scored live from your real memory, thresholds printed with the numbers. Any memory stack a connector can scan could be graded by the same exam.

`helicon fleet` is **one screen for the whole fleet**, and it is project-first because that is the unit a person thinks in. Every field is DERIVED and none is typed: where a project *stands* comes from git (branch, HEAD, uncommitted, unpushed, commits in 24h), *needs you* from runs that finished with no verdict, *friction* from the complaint log, *next* only from a prompt you actually accepted **and** that matches this project — otherwise it says `unmeasured` and names what would fill it. A hand-typed "next step" field is stale in a week, and then you have rebuilt the dashboard problem one layer down.

It opens with IDLE, because that is the only line that costs money while you read it: terminal-hours sitting silent right now. Reported as a **floor, not a total** — a session thinking hard writes nothing and reads as idle — and counted only for sessions a human actually typed into. The first version counted every transcript and claimed 241 idle hours across 74 "terminals"; 68 were headless judge processes that had already exited.

It ends with **YOU CAN**, which is the section that is not about state. On the day this was built, eight terminals sat idle for two hours while prompts were hand-written for a human to carry between them, because no agent instance remembered that `ListAgents` and `SendMessage` address every peer session directly. A capability nobody recalls is not a capability, so the screen states it: an agent should leave this screen holding an ability it walked in without.

`helicon complaints` is **the complaint log**: every time you pushed back on an agent, recovered from your own transcripts and grouped by kind. Every other signal in this repo is the machine grading itself — battery verdicts, judge runs, self-scored retrieval. A correction is a human saying *no, that is wrong* with nothing to gain by lying, and it is destroyed the moment the terminal closes.

It rests on two authorship gates, both measured before they were written. A `type: user` entry is **not** necessarily the human: across ~600 local transcripts, 46% of non-tool user turns were programmatic judge runs, task notifications, messages from *other* agent sessions, or injected skill files. Count those as feedback and an agent is grading itself on its own prose. And even a turn you typed is not always your writing — pasted lane prompts are long and any correction-shaped phrase inside them sits deep in the body, so position separates them. Raw user turns score 28% precision; both gates ~81% (self-graded, the author read all 46). Storage is a cube of `type='complaint'`, so it inherits FTS, decay and the review loop — no new table.

On this repo's own store the largest category is not a wrong fact. It is **wrong-plan**: the agent choosing the wrong next thing to do.

## Audit a store you don't own

The exam is not limited to your own memory. Any repo with a committed agent-rules file (AGENTS.md, CLAUDE.md, .cursorrules, ...) is a memory store someone's agent obeys every session — so it can be examined:

```bash
bash scripts/demo_public_store.sh          # default: openai/codex AGENTS.md
```

This replays the file across its REAL git history (no staging): ingests an old commit, snapshots retrieval, replays to HEAD, reconciles, runs the rot exam. On openai/codex (27 real commits of AGENTS.md edits, cited by SHA in the output): 5 sections retired as drifted, 1/1 retrieval snapshot regressed — R10 and R8, live, on a store we don't own. Reproducible by anyone.

## The life-OS benchmark — scored against human-labeled rot

On Jul 5 a 5-agent manual audit swept the operator's real second brain (Obsidian vault + Claude Code memory dir), archived 33 stale docs and stamped 21 drifting docs with dated `> **LOUPE` correction banners. Those banners are a labeled dataset of real memory rot. The benchmark ingests the same corpora with the banners stripped (the answer key never enters the input) and scores the deterministic detectors against them:

```bash
python3 scripts/rot_bench_lifeos.py    # read-only on sources, throwaway DB, zero LLM
```

Honest numbers from the first run (232 files, 1,667 section memories): **6/16 file-level catches, 4/16 strict facet-match** — the output labels the difference itself. What it caught: both merge-status flips (audit doc still said 'NOT patched' after the fix merged), a stale dashboard doc, a dead 7-week-old plan. What it found that the humans missed: a win-count fight (9 vs 10) living in the resume and two application drafts, and 35 files still asserting a dead project name post-rebrand. Named misses, on the roadmap: overlapping-date-range drift (Aug 14-22 vs Aug 15-24 overlap, so interval semantics reads agreement), living-doc supersession without a declared rename, and content-based staleness (a young file asserting old facts).

## Access & trust model (read this before connecting your vault)

A tool that audits your memory reads your memory. That access is scary, so here is exactly what Mountain of Helicon does with it — from the code, not a promise:

**Reads (always read-only):** your configured sources — Claude Code transcripts, Obsidian vault, git repos, rules files, memory stores via adapters. Connectors never write to a source. The life-OS benchmark and the rot exam open the store read-only.

**Writes, exhaustively:**
- its own SQLite DB and `data/` (findings, verdicts, drift reports, compiled context)
- `helicon fix-skills` and other write-backs: **dry-run by default**, `--apply` required, `.bak` written next to every file before modification, second run is a no-op
- `helicon watch --install`: one tagged line in your crontab, removed by `--uninstall`
- `helicon compile`: compiled context files under `data/compiled/` (`--output` redirects). It writes nothing into `~/.claude/`: an auto-inject path exists in the source (`compiler.inject_into_claude_code`) but no command calls it, so the pull path (`helicon_context` over MCP) is the working half of that loop and the push half is unwired. Our own store still carries pre-rename `glaze-*` skill files from an older injector, which the skills audit flags: a live example of why write-backs need lifecycle discipline
- `helicon policy --inject` (alias `helicon gold --inject`): `~/.claude/GOLDEN_RULES.md`, dry-run by default, `.bak` kept
- your vault: **never**. Corrections are memories in Helicon's store, not edits to your files. You stay the only writer of your second brain.

**Leaves your machine:** nothing, unless you configure a Qwen key — then excerpts of candidate memories (truncated content) go to the model for judging, and the response is cached locally. Keyless mode runs every deterministic check with zero egress and says so instead of degrading silently.

**Decisions:** every destructive or state-changing action (kill, retire, resolve, dismiss, rule application) is either made by you or made by a written rule you previewed and approved — and automated decisions are quarantined from the learning loop (rot class R9), so the tool cannot launder its own output into your evidence.
## Your domain, your lexicon (config, not code)

The claim-conflict detectors ship with built-ins (win counts, episode numbers, merge status, decision status) and take the rest from `config.json` — an enterprise wiki or research vault declares its own counted things and polar statuses, and gets the same conflict machinery, evidence receipts and resolve loop:

```json
"claims": {
  "metrics":   {"headcount": "\\b(\\d{2,5})\\s+employees\\b"},
  "statuses":  {"contract": {"live": "contract (is )?live", "expired": "contract (is )?expired"}},
  "canonical": {"wins": "mindmap.md"}
}
```

`canonical` encodes the single-source-of-truth rule: declare WHERE a fact's truth lives, and a conflict files as *"Drift from canon: canon says 9; 8, 10 asserted elsewhere"* — the human confirms a pre-decided direction instead of adjudicating from scratch.

Doc honesty is enforced: `python3 -m helicon.docdrift` compares this README's numeric claims against counts computed from source, and it runs in the test suite — stale docs fail the build. (It caught this very README claiming 20 commands the hour the 21st landed.)

Everything destructive is dry-run by default and takes `--apply`.

## Honest eval numbers

**None of the four numbers below reproduces on a fresh clone, and that is a gap, not
a footnote.** They were computed against the author's own populated store — roughly
6,900 memories carrying real human kill/approve rulings — which is personal data and
is not in this repository. On a clean install `helicon eval` has no store to read and
says so rather than printing anything. Treat these as *measured once, on a corpus you
cannot see*, which is weaker evidence than the doorway numbers above, every one of
which you can reproduce on your own repo in seconds.

- Composite: **~67** (live, as of 2026-07-13 — run `helicon eval` to recompute; retrieval P@3 + MRR + decay-AUC; audit axis excluded -- no labeled ground truth).
- Retrieval: P@3 0.615, MRR 0.596. Small internal benchmark (n=13, one label per query) -- disclosed, not hidden.
- **Decay predicts human kills at rank-AUC 0.78** (mean confidence of killed memories 0.14 vs approved 0.27). A real, independent signal.
- Consolidation: ~9-10x fewer tokens; Qwen-judged quality favors synthesis (self-graded, shown as direction, not proof).
- The public demo store is 19 labelled planted memories and contains no personal
  data. Live scans read only the sources each user configures.

## Built on established patterns, extended

Mountain of Helicon's capabilities stand on well-understood memory-systems patterns and take each one further. The lineage, stated honestly — the second column is the established idea, the third is our own build on top of it:

| Capability | Established pattern | How Mountain of Helicon extends it |
|-----------|--------|---------------------------|
| Versioned memory units | Structured memory units, not raw text | HeliconCube: source, hash, valid_from, confidence, decay per type |
| Multi-axis audit | Temporal/factual/logical consistency checks | 13-class rot exam, each with a receipt and a never-twice guard |
| Weibull decay | Non-uniform forgetting curves | Per-type kappa, and decay rank-predicts human kills (AUC 0.78) |
| Novelty gate | ADD/NOOP/MERGE at ingestion | Gate + provenance, so a merge never loses the source it came from |
| Anti-confabulation | Challenge claims against evidence | Grounding check + R12 phantom-association catch |
| Retrieval learning | Track surfaced vs acted-on | Q-value ranking rewarded by human rulings only — no self-echo |
| Identity & phantom coherence | *(ours — no store or prior system does this)* | R11 fork detection, R12 phantom catch, rulings compiled to law |

## Architecture

<p align="center">
  <img src="docs/architecture.svg" alt="Mountain of Helicon architecture: the store, retrieve, output, attribute, rule, law loop" width="100%">
</p>


- **Backend:** Python 3.12, FastAPI (154 endpoints), SQLite + FTS5 (43 tables declared across the memory and correction stores). **Qwen-native retrieval when a Model Studio key is configured**: `text-embedding-v4` (1024-dim) dense vectors + FTS5, fused by Reciprocal Rank Fusion, then a `qwen3-rerank` two-stage pass, the whole retrieve→rerank stack on Alibaba Cloud (falls back to local MiniLM + linear fusion, FTS-only, when no key)
- **Frontend (optional):** React 19, TypeScript, Vite. Four surfaces — **Next Moves** (memory state → cited next prompts/goals, generated by Qwen, every move citing the memory it came from), **Memory** (sources, review coverage, health), **Needs Ruling** (every failed check with why/evidence/action, grouped Drift / Stale / Smartness), **Golden Rules** (rulings compiled with provenance, injectable). The dashboard is one of three interfaces (CLI · MCP-in-IDE · dashboard)
- **AI:** Qwen Cloud API via OpenAI-compatible SDK (see table above)
- **Distribution:** BYOK + local-first. No hosted personal-store service is
  advertised for v0.1; public hosting waits for HTTPS, sessions, configured
  CORS, rate limits, and backups.

## License

MIT
