Metadata-Version: 2.4
Name: mnemekit
Version: 0.5.0
Summary: A zero-LLM, zero-embedding, cue-indexed tag memory for conversational agents.
Author: Yam
License-Expression: MIT
Project-URL: Homepage, https://github.com/FTP2026/Mneme
Keywords: memory,llm,agent,conversational,retrieval,spreading-activation
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: spacy>=3.7
Provides-Extra: zh
Requires-Dist: jieba; extra == "zh"
Provides-Extra: eval
Requires-Dist: mem0ai; extra == "eval"
Requires-Dist: sentence-transformers; extra == "eval"
Requires-Dist: matplotlib; extra == "eval"
Requires-Dist: pararun; extra == "eval"

# 🧠 Mneme

**A cue-indexed tag memory for conversational agents — no embeddings, no LLM in the memory path.**

[![PyPI](https://img.shields.io/pypi/v/mnemekit.svg)](https://pypi.org/project/mnemekit/)
[![Python](https://img.shields.io/pypi/pyversions/mnemekit.svg)](https://pypi.org/project/mnemekit/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)

Most agent memories embed every turn and distill it with a write-time LLM, then search by vector similarity — expensive and opaque. Mneme instead does what human recall does: **index by cues, reinstate by cue-match, and spread activation along both shared meaning and the thread of conversation** — with a plain inverted index over deterministic tags. No vector database, no language model, no embeddings.

> **At zero memory cost, under an identical controlled comparison, on each benchmark's primary metric Mneme is on par with a strong dense retriever** (LoCoMo F1 52.4±0.1 vs 52.9; LongMemEval acc 62.8 vs 62.2) **and beats the cheaper baselines** (mem0 and BM25 on LoCoMo, BM25 on LongMemEval). Same answerer, judge, prompt, and budget; only the memory differs.

---

## ✨ Why Mneme

- **Zero cost.** No write-time LLM calls, no embeddings, sub-millisecond recall, storage is plain JSON files.
- **Competitive quality.** Matches or beats embedding + LLM memory on standard benchmarks (see below).
- **Interpretable.** Retrieval is explicit cue overlap ranked by IDF — no opaque similarity threshold to tune.
- **Human-like forgetting.** What has no distinctive cue, or whose cues never recur, is simply never surfaced — forgetting as retrieval failure, for free.
- **Local & sovereign.** Your memory lives in a local transactional SQLite store. English by default; Chinese optional.

## 📊 Benchmarks (controlled: same answerer/judge/prompt/budget, only the memory differs)

**LoCoMo** — low-distractor dialogue (F1 / accuracy):

| method | cost | F1 | acc |
|---|---|---|---|
| mem0 (LLM extract + vectors) | LLM+emb | 40.4 | 57.3 |
| BM25 (raw turns) | zero | 45.8 | 57.0 |
| **Mneme** (hybrid edges) | zero | **52.4** | **67.5** |
| dense retriever (bge-large) | embeddings | 52.9 | 67.7 |

**LongMemEval** — high-distractor, ~50 sessions per question (accuracy):

| method | cost | acc |
|---|---|---|
| BM25 (raw turns) | zero | 61.2 |
| dense retriever (bge-large) | embeddings | 62.2 |
| **Mneme** (hybrid) | zero | **62.8** |

Zero-cost Mneme **is on par with dense retrieval on both benchmarks**, at no embedding or write-time cost, and it is reader-agnostic (near-identical with two different answerer LLMs) and interpretable. Which edge type helps is regime-sensitive: discourse edges clearly help clean dialogue (~5 points on LoCoMo), while on high-distractor histories the edge configurations differ only within run-to-run noise.

## 🚀 Install

```bash
pip install mnemekit
python -m spacy download en_core_web_sm     # English model (required)

pip install "mnemekit[zh]"                   # optional: Chinese (jieba)
```

### New in 0.4

Mneme now separates access, evidence, and immutable source provenance. The
canonical search projects the activated tag/event graph into deduplicated,
role-preserving evidence containers under an explicit item and token budget.
Typed causal provenance participates inside the same conserved activation field;
there is still no embedding, language model, or learned reranker in the memory
path. The existing `Memory.recall` interface and old stored turns remain
compatible.

### New in 0.5

Mneme now persists turns and events incrementally in a transactional SQLite WAL
store. Compressed retrieval-state checkpoints make reopening proportional to a
short tail rather than the full turn history. `remember_many()` commits an
ordered multi-message Add atomically, and 0.4.x JSON stores remain readable with
an explicit migration path for subsequent writes.

## ⚡ Quick start

```python
from mneme import Memory

m = Memory(root="~/.mneme")

m.add(
    "My dog Lucky is a golden retriever",
    role="user",
    session_id="chat-42",
    round_id=1,
    exchange_id="chat-42:x1",
)
m.add(
    "Cute! How old is Lucky?",
    role="assistant",
    session_id="chat-42",
    round_id=2,
    exchange_id="chat-42:x1",
)

m.search("what breed is my dog")     # -> [SourceTrace(...), ...]
m.search_evidence(
    "what breed is my dog",
    topn=50,
)                                      # -> [EvidenceItem(...), ...]
```

`add` preserves every role as a separate immutable source trace. Shared exchange
and event identifiers organize traces without concatenating them. `search`
activates the stored network once, takes one global Top-K, and reinstates the raw
source text. The original `remember(query, response)` and `recall` interfaces
remain available for compatibility.

`search_evidence` separates access from returned evidence. The default
`TAG_GRAPH_DEDUP` projection puts direct query tags and tags reached through graph
activation in one evidence pool. Each tag is one result item containing the raw
source spans attached to that cue. Typed response provenance can reinstate its
causal source inside the same activation field; it does not append context after
selection. Spans are ordered by activation, exact activation ties prefer the
shorter lossless source, and a source is emitted only once across items. Every
span retains its source ID, role, timestamp, and full text. Explicit `TAG` and
`TAG_GRAPH` preserve the corresponding non-deduplicated views for ablations.
Top-K is applied to evidence items after projection; it is never followed by a
hidden expansion step. The library default is 50 items; callers may request the
Agent Memory Leaderboard maximum of 100.

By default, an exchange is also its lossless micro-event. Pass the same explicit
`event_id` to organize several exchanges into one evolving event. Mneme does not
silently merge exchanges using similarity thresholds or timing rules.

For ordered multi-message ingestion, `remember_many()` accepts a list of
`RememberInput` values and commits them in one SQLite transaction. A later search
sees the complete batch immediately; a failed batch exposes none of its turns.

```python
from mneme import Memory, RememberInput

m.remember_many([
    RememberInput(query="I moved to Kyoto.", session_id="chat-42", round_id=3),
    RememberInput(response="Noted.", session_id="chat-42", round_id=4),
])
```

New stores use `store.sqlite3` in WAL mode and periodically checkpoint the full
derived retrieval state. Reopening loads the latest checkpoint and only replays
its short tail. A 0.4.x JSON store remains readable but is write-protected until
the caller explicitly invokes `migrate_store()`:

```python
m = Memory(root="path/to/legacy-store")
m.migrate_store()
```

The synthetic capacity benchmark is available as
`python benchmarks/store.py --checkpoints 100 1000 5000 10000`. Use
`--batch-size 20` to measure the batch path used by multi-message Add requests.

## 🔍 How it works

```
INGEST (no LLM, no embedding)
  turn ──▶ tag extraction (spaCy NER + lemmas / jieba) ──▶ inverted index  (tag → turns)
                                                           + event tag inheritance

RECALL (per query) — spread over an ephemeral heterogeneous graph, then discard it
  query ─▶ ① lexical: IDF cue reinstatement ─seeds─┬─▶ ② associative edge (shared tag, 1 hop)
                                                   └─▶ ③ discourse edge (adjacent turn in event)
                                                                        │
                                                          recalled context ─▶ your reader
```

1. **Cue index.** Each turn is tagged deterministically; tags form an inverted index (a sparse cue index, à la hippocampal indexing). Keyword-free follow-ups inherit their event's tags, so they stay recallable.
2. **Reinstatement, then two kinds of edge.** A query reinstates directly-cued turns lexically, then spreads activation over an ephemeral graph with two edge types: **associative** edges (shared discriminative cue → one hop) reach the *associative tail* a direct cue can't; **discourse** edges (adjacent turns in the same event) pull in the local conversational context around a hit. Neither is materialized — both are read straight off the index and event structure, used for the one query, and discarded.

## 🀄 Chinese & custom dictionaries

```python
m = Memory(root="~/.mneme", userdict="my_terms.txt")   # jieba user dictionary
m.remember("我养了只金毛狗，叫 Lucky", "可爱！", session_id="u1", round_id=1)
```

Chinese is segmented by `jieba`; a custom dictionary keeps vertical-domain terms intact.

## 📄 Paper & reproduction

The write-up (method, cognitive grounding, controlled comparison, ablations) is in [`paper/`](paper/); the evaluation harness (LoCoMo / LongMemEval / mem0) is in [`eval/`](eval/). The frozen reproduction version is on the `paper-repro` branch.

## 📌 Citation

```bibtex
@misc{mneme2026,
  title  = {Do Conversational Agents Need Vectors? A Zero-LLM, Zero-Embedding Tag Memory},
  author = {Hao, Shaochun},
  year   = {2026},
  note   = {https://github.com/FTP2026/Mneme}
}
```

## License

MIT
