Metadata-Version: 2.4
Name: stratus-engine
Version: 0.2.0
Summary: Two-layer memory infrastructure for LLM applications.
Keywords: llm,memory,rag,vector-database,agents
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Requires-Python: >=3.12
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pydantic>=2.7
Provides-Extra: mongo
Requires-Dist: pymongo>=4.8; extra == "mongo"
Provides-Extra: vector
Requires-Dist: chromadb>=0.5.0; extra == "vector"
Provides-Extra: openai
Requires-Dist: openai>=1.40.0; extra == "openai"
Requires-Dist: tiktoken>=0.7.0; extra == "openai"
Requires-Dist: python-dotenv>=1.0; extra == "openai"
Provides-Extra: anthropic
Requires-Dist: anthropic>=0.40.0; extra == "anthropic"
Provides-Extra: all
Requires-Dist: anthropic>=0.40.0; extra == "all"
Requires-Dist: chromadb>=0.5.0; extra == "all"
Requires-Dist: openai>=1.40.0; extra == "all"
Requires-Dist: pymongo>=4.8; extra == "all"
Requires-Dist: python-dotenv>=1.0; extra == "all"
Requires-Dist: tiktoken>=0.7.0; extra == "all"
Provides-Extra: demo
Requires-Dist: chromadb>=0.5.0; extra == "demo"
Requires-Dist: openai>=1.40.0; extra == "demo"
Requires-Dist: pymongo>=4.8; extra == "demo"
Requires-Dist: pytest>=8.0; extra == "demo"
Requires-Dist: python-dotenv>=1.0; extra == "demo"
Requires-Dist: tiktoken>=0.7.0; extra == "demo"
Provides-Extra: dev
Requires-Dist: build>=1.2; extra == "dev"
Requires-Dist: pytest>=8.0; extra == "dev"
Requires-Dist: twine>=5.0; extra == "dev"
Dynamic: license-file

# Stratus Engine

Stratus Engine is a pragmatic memory layer for LLM applications. The MVP keeps
active conversations in MongoDB, promotes only useful facts into long-term
memory in Chroma, and warms relevant context when a session is reopened.

The point is simple: do not make vector search behave like a chat log. Keep the
active session close, then send only valuable extracted memories to the long-term
vector store.

## Install and use

Install the minimal, dependency-free core:

```powershell
pip install stratus-engine
```

Or install a backend combination from source:

```powershell
pip install -e ".[mongo,vector,openai]"
```

The public API starts with one factory:

```python
from stratus_engine import StratusConfig, create_engine

engine = create_engine()  # local, in-memory, no API key or Docker needed
session = engine.create_session("user_123", title="Product assistant")

engine.append_user_message(session.id, "I prefer React for dashboards.")
engine.extract_memories(session.id)

context = engine.build_context(session.id, "What frontend preference do I have?")
print(context.as_prompt_sections())
```

Switch to persisted sessions and Chroma vector memory without changing the
application flow:

```python
engine = create_engine(StratusConfig(
    session_backend="mongo",
    memory_backend="chroma",
    use_openai=True,
))
```

For advanced setup, import the storage and provider adapters directly. The
factory is the stable plug-and-play path; the internal modules are intentionally
more configurable and may evolve faster.

## The Private Kitchen Analogy

Think of the LLM as a chef serving a customer. The temporary memory layer is
the kitchen: active conversation context and likely-useful memories are kept
close for fast service. The permanent vector database is the warehouse: it
stores durable facts that are available when needed.

When a session reopens, Stratus prepares the kitchen by running intelligent
warmup queries and moving relevant memories into the temporary layer. If a new
request needs something that is not already there, the engine makes an ad-hoc
trip to permanent memory—like fetching a special ingredient from the
warehouse. This keeps vector search useful without making it the default path
for every message.

## Current MVP

- MongoDB session layer for active conversations
- Structured memory extraction from user messages
- Chroma vector DB adapter for long-term memories
- Optional OpenAI embeddings and live answer generation
- Session reopen cache warming with predefined queries
- Context assembly that prefers active session state before ad-hoc recall
- Metrics for estimated context tokens, retrieved memories, and relevance
- Benchmark script for vector-search reduction
- Runnable local demo and tests

## Quick Start

```powershell
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -e ".[dev]"
python examples/demo.py
python examples/benchmark.py
pytest
```

For an LLM-powered demo, copy `.env.example` to `.env`, add your
`OPENAI_API_KEY`, and run `python examples/live_openai_demo.py`. The offline
tests and local demo do not require an API key.

The default warmup planner is deterministic and derives retrieval intents from
the session. You can opt into `OpenAIWarmupPlanner` when you want the model to
produce more nuanced, structured retrieval plans; the engine keeps the
deterministic planner as a safe fallback.

## Model providers

The engine is provider-agnostic. The memory pipeline depends on small Python
protocols, while model and embedding clients are optional adapters.

```python
from stratus_engine import OllamaConversationClient

llm = OllamaConversationClient(model="llama3.2")
result = llm.respond(prompt="What do you remember?", context=context)
```

Available adapters include OpenAI, Anthropic Claude, and local Ollama. Ollama
requires a running Ollama server and a pulled model; Claude requires
`pip install -e ".[anthropic]"`. The deterministic local mode remains useful
for tests and demos without any model provider.

## Package layout

```text
stratus_engine/
├── core/       public engine and domain models
├── storage/    MongoDB, Chroma, and in-memory adapters
├── providers/  model and embedding integrations
├── retrieval/  candidate ranking and deduplication
├── extraction/ memory extraction strategies
└── planning/   dynamic session warmup planning
```

Chroma retrieval is intentionally staged: it fetches a wider candidate pool,
combines vector similarity with lexical coverage and memory quality, applies
recency/access signals, removes near-duplicate facts, and only then returns the
final context. This is durable-memory retrieval, not a naive every-message RAG
loop.

## Evaluation and comparison

Run the reproducible local comparison:

```powershell
python examples/evaluate_architectures.py
python examples/end_to_end_simulation.py
```

The evaluation compares three explicit strategies:

| Strategy | What it does | Main tradeoff |
| --- | --- | --- |
| `full_context` | Sends every durable fact on every turn | High context cost, no retrieval calls |
| `naive_retrieval` | Runs retrieval for every turn | More retrieval calls and noisy context |
| `hybrid_warmup` | Warms likely facts once and gates ad-hoc recall | More orchestration, lower repeated retrieval |

The pack reports recall, vector-call count, and estimated context tokens. It is
intentionally small and inspectable so contributors can understand every case.
It is not presented as a universal benchmark. Production evaluations should add
domain-specific conversations, adversarial queries, stale facts, conflicting
facts, multilingual data, and human or LLM-judged answer quality.

The end-to-end simulation shows the complete lifecycle offline: session writes,
durable memory promotion, engine recreation, login-time warmup, continued work,
retrieval traces, and context analytics. Its traced in-memory store can be
replaced with `ChromaLongTermMemoryStore` without changing the engine flow.

For a live interactive version, run:

```powershell
python examples/live_session_simulation.py --provider openai
python examples/live_session_simulation.py --provider ollama --model llama3.2
python examples/live_session_simulation.py --provider offline
python examples/live_session_simulation.py --provider openai --backend chroma
```

The live simulation recreates the engine after seed memories are promoted, then
accepts your questions and prints the actual model response plus retrieval
reason, warmed memories, recalled memories, estimated tokens, and relevance.
Use `--backend chroma` with the Docker Chroma service to exercise a real vector
database; the default `memory` backend is faster for repeatable local demos.

## Docker Services

Start MongoDB and Chroma:

```powershell
docker compose up -d mongo chroma
```

## Chroma Vector Demo

This uses MongoDB for sessions and Chroma for long-term memory. It uses local
hash embeddings, so it does not need an OpenAI key.

```powershell
pip install -e ".[mongo,vector,dev]"
python examples/chroma_demo.py
```

## Live OpenAI Demo

This uses MongoDB, Chroma, OpenAI embeddings, and an OpenAI chat completion. It
prints context metrics and API token usage.

Create a local `.env` file with `OPENAI_API_KEY` set to your key. The live demo
loads that file automatically, or you can set the variable in the current
PowerShell session before running it.

```powershell
pip install -e ".[demo]"
$env:OPENAI_API_KEY = "your_api_key_here"
python examples/live_openai_demo.py
```

## Benchmark

The benchmark compares naive retrieval, where every prompt would run vector
search, against Stratus's gated recall.

```powershell
python examples/benchmark.py
```

It prints vector-search reduction, average estimated context tokens, average
memory relevance, and a prompt-level trace.

## Programmatic Usage

```python
from stratus_engine import (
    ChromaLongTermMemoryStore,
    MongoSessionStore,
    OpenAIEmbeddingProvider,
    StratusEngine,
    analyze_context,
)

engine = StratusEngine(
    session_store=MongoSessionStore("mongodb://localhost:27017"),
    memory_store=ChromaLongTermMemoryStore(
        host="localhost",
        port=8000,
        embedding_provider=OpenAIEmbeddingProvider(),
    ),
)

session = engine.create_session("user_123", title="Demo")
engine.append_user_message(session.id, "I prefer React over Angular.")
engine.extract_memories(session.id)
engine.reopen_session(session.id)

context = engine.build_context(session.id, "What frontend preference do I have?")
metrics = analyze_context("What frontend preference do I have?", context)
print(context.as_prompt_sections())
print(metrics)
```

## Demo Story

The demo shows five important behaviors:

1. Active conversation context is read from the session layer.
2. Useful facts are extracted into structured memories.
3. Chroma retrieves long-term memories.
4. Reopening a session warms relevant context before the user asks anything.
5. Metrics show token footprint and retrieval relevance.

## Design Direction

Next steps are intentionally narrow: replace the heuristic extractor with an
LLM-backed extractor, add background extraction, and add a small evaluation set
that proves latency, recall quality, and reduced vector-search usage.
