Metadata-Version: 2.4
Name: sangam
Version: 0.1.18
Summary: Intelligent type-checking based code analysis and generation CLI tool
Author: sangam team
License: MIT
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Code Generators
Classifier: Topic :: Software Development :: Quality Assurance
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Requires-Dist: mypy>=0.900
Requires-Dist: flask>=2.0
Dynamic: author
Dynamic: classifier
Dynamic: description
Dynamic: description-content-type
Dynamic: license
Dynamic: requires-dist
Dynamic: requires-python
Dynamic: summary

# sangam

> An **LLM harness** — a runtime and orchestration framework for coding agents
> backed by large language models.

`sangam` is being built as an **LLM harness**: a runtime that hosts
LLM-driven coding agents, gives them isolated places to run code, and a typed set
of tools to call. The design rests on four pieces —

- **Runtime-agnostic execution** — a single `Executor` interface
  (`RunSpec` in, `RunResult` out) with pluggable backends (Python venv, Node/nvm,
  Docker/Podman, subprocess) so the agent never depends on a specific runtime.
- **Typed tool registry** — each tool (read/write/edit files, run code,
  type-check, run tests, search) exposes a JSON schema, letting the LLM act
  through function-calling.
- **Model client** — an LLM abstraction (`complete` / `chat` / `tools`) that can
  talk to any provider: any OpenAI-compatible endpoint, or local models (Ollama /
  LM Studio). When no API key is set it falls back to a local model automatically.
- **Agent loop** — plan → act (call a tool) → observe (tests/lint/output) →
  reflect, with conversation memory and resumable state.

It started as a mypy-focused natural-language CLI, and those capabilities — mypy
type analysis, file operations, and resumable task history — stay on as a
built-in toolset and a working surface. Full architecture in
[`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md).

## Status

The harness is under active construction: the core agent loop, multi-runtime
executors, tool registry, model client, session store, headless HTTP API server,
VS Code extension, and the local webapp are in place (Phases 1–4 and 6
complete). The desktop app surface is the remaining roadmap. See
[`todo.md`](todo.md) for the roadmap.

**Available now**

- Core engine split from the CLI surface (`core/engine.py`) — usable as a library.
- `Executor` protocol + `RunSpec` / `RunResult` runtime contracts
  (`core/runtime/base.py`).
- **Multi-runtime executors** — pluggable backends behind the single `Executor`
  interface:
  - `VenvExecutor` (`core/runtime/venv.py`) — creates an isolated venv, installs
    deps via pip, runs Python, captures stdout/stderr/exit.
  - `NvmExecutor` (`core/runtime/nvm.py`) — runs JS/React via nvm + npm
    (`nvm use`, `npm install`, run/build).
  - `ShellExecutor` (`core/runtime/shell.py`) — runs native bash/zsh scripts and
    one-off commands, auto-detecting the user's shell.
  - `ContainerExecutor` (`core/runtime/container.py`) — runs arbitrary code in
    Docker or Podman (image pull, run).
  - `SubprocessExecutor` (`core/runtime/subprocess.py`) — plain subprocess
    fallback backend.
- **Runtime auto-detection** (`core/runtime/detect.py`) — maps a runtime name
  (`python`/`venv`, `node`/`js`/`react`/`npm`, `shell`/`bash`/`zsh`,
  `container`/`docker`/`podman`, `subprocess`/`direct`) to the matching backend.
- **Environment registry** (`core/runtime/registry.py`) — records created venvs
  and npm projects in `~/.sangam/environments.json` so they are reused across
  runs instead of recreated.
- **Sandbox Manager** (`core/sandbox.py`) — `SandboxPolicy` + `SandboxedExecutor`
  wrapping any backend with filesystem scope (path-level validation, symlink
  resolution), wall-clock timeouts (default/max clamping), network isolation
  (proxy-env clearing), and POSIX resource limits (`cpu_time`, `fsize`,
  `nofile`) via `resource_limit_preexec`.
- **Agent loop** — `Agent` (`core/agent.py`) orchestrating `Planner`
  (`core/planner.py`), `Memory` (`core/memory.py`), the `ToolRegistry`, and an
  optional `ModelClient` in a plan → act → observe → reflect loop. Runs in a
  deterministic mode when no model client is configured.
- **Model client** (`core/model/`) — `ModelClient` protocol with
  `OpenAIModelClient` (any OpenAI-compatible endpoint) and `LocalModelClient`
  (local providers such as Ollama / LM Studio); credentials read from env/config.
  `build_model_client(config)` picks the backend — the entry selected from the
  model catalog when one is configured, otherwise OpenAI when an API key is set,
  otherwise a local model (auto-detecting an available model from the running
  server when `MODEL_NAME` is unset).
- **Model catalog** (`core/model/catalog.py`) — `~/.sangam/models.yaml` is the
  source of truth for which models sangam can use, following Continue's
  `config.yaml` `models` shape (`name`, `provider`, `model`, optional `apiBase` /
  `apiKey` / `roles` / `capabilities` / `defaultCompletionOptions` /
  `requestOptions`). The active chat/agent model is the file's `default` entry,
  else the first entry with role `chat`, else the first entry. `sangam models`
  lists the configured entries merged with the ids advertised by probed local
  runtimes (Ollama, LM Studio, and every configured `apiBase`) via `/v1/models`.
  sangam never installs a model runtime.
- **Typed tool registry** (`core/tools/`) — `read_file`, `write_file`,
  `edit_file`, `list_files`, `search`, `run_python`, `run_js`, `run_container`,
  `run_shell`, `install_deps`, `type_check`, `run_tests`, `suggest_fixes`,
  `web_search` (web search + page fetch), each with a JSON schema for LLM
  function-calling.
- **Session store** (`core/session_store.py`) — SQLite-backed sessions, messages,
  model/token usage, and FTS5 search (`~/.sangam/sessions.db` by default).
- **CLI subcommands** — `sangam agent "<goal>"`, `sangam models`,
  `sangam exec --runtime <rt> <command>`, `sangam serve` (headless HTTP API
  server), `sangam web` (local webapp UI + same-origin API proxy), and
  `sangam session` (list stored chat sessions + generate their descriptions),
  `sangam restore` (continue a stored chat session in interactive mode),
  `sangam run` (one-shot prompt through the interactive chat pipeline),
  plus the legacy single-task positional, REPL, and batch mode.
- **Headless HTTP API server** (`sangam/server/`, Flask) — exposes the agent,
  tools, runtimes, and session store over local HTTP with SSE streaming; the
  single backend for the VS Code extension, local webapp, and desktop app.
- **Local webapp** (`webapp/` + `sangam web`) — a single-page UI (chat, file
  tree, diagnostics, agent-step viewer, palette, status bar) served on
  localhost with the API proxied on the same origin — no build step, no CDN.
  Its presentation layer is built on the DeepSeek Harness design system
  (three-layer `--sgw-*` tokens with light/dark themes, the `AppFrame` shell
  geometry) and it renders markdown, syntax-highlighted code, unified diffs
  and collapsed reasoning rows.
- **`install.sh`** — pulls sangam from GitHub `main`, installs it into a venv at
  `~/.sangam/venv`, and runs the headless server. Before starting it probes for a
  local model runtime (Ollama / LM Studio); if none is reachable it errors out
  with "no model available/selected" rather than installing one.
- mypy integration wired through the Executor backends — analyze a file or a
  whole directory, JSON output parsing, and fix suggestions. Defaults to
  `SubprocessExecutor` (`sys.executable -m mypy`); opt-in `use_venv=True` →
  `VenvExecutor` with mypy installed; custom backend via `executor=`.
- Regex-based task engine — natural-language commands mapped to file operations
  and mypy analysis.
- File operations — read, write, append, delete, list, and extract
  functions/classes/imports.
- JSONL task history with search.
- Interactive REPL and batch mode.

**Roadmap**

- VS Code extension polish — chat panel, diagnostics, inline actions (extension
  scaffolded and installable; UI activation pending a reload).
- Cross-platform desktop app — React UI over the Python core behind a native
  shell, packaged for macOS (`.app`/`.dmg`) and Windows (`.exe`/`.msi`); reuses
  the headless server as its backend.
- Hardening — `doctor` capability probing, resumable `history`, runtime/model/
  sandbox config, end-to-end sandbox limits, CI-ready non-interactive mode.

## Local Webapp

The **local webapp** is a zero-install surface for the same harness: five
static files (`index.html`, `app.css`, `app.js`, `markdown.js`, `diff.js`) —
no build step, no bundler, no CDN, no runtime dependency — served on localhost
by a small Python static server, talking to the core over HTTP. It reuses the
headless HTTP API server from Phase 4 as its backend — another thin surface
over the same core, not a separate engine.

Item 6.8 restyled its presentation layer onto the design system and app-shell
geometry of [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)
(tokens re-prefixed `--sgw-*`; attribution and the full token/geometry tables in
[`docs/features/webapp-dsh-restyle.md`](docs/features/webapp-dsh-restyle.md)).
No DeepSeek branding is used and no DSH asset is redistributed — only the
design system, re-implemented in plain CSS.

```bash
sangam web                    # UI on http://127.0.0.1:8899, API proxied on one origin
sangam web --port 9000        # different UI port (API stays on 8765)
sangam web --no-browser       # don't open the browser
sangam web --webapp-dir ./webapp   # explicit webapp directory (or $SANGAM_WEBAPP_DIR)
sangam web --theme dark       # pin the page's theme (or $SANGAM_WEB_THEME)
```

The page follows your **system** theme by default and repaints when the OS
changes it — no reload. Use the ☀ / ☾ / ⌘ switch in the top bar to override;
the choice is remembered in the browser. `--theme` pins it from the command
line for a fresh profile, a kiosk or a screen share, where there is no stored
value and no OS signal to trust. See [Theming](#theming).

What you get in the browser (mirroring the sangam desktop layout):

- **Chat** — send goals in **Agent** mode: the plan → act → observe loop
  streams live from `/api/agent/stream` (step cards for every tool action and
  observation, the final answer, and a stop button). **Exec**, **Type check**,
  and **Tests** composer modes map to `/api/exec`, `/api/type_check`, and
  `/api/run_tests` for one-shot tool runs (per-mode hints + tooltips explain
  what your text becomes and which endpoint it hits).
- **Model selector** — the composer's model chip lists the catalog from
  `GET /api/models` (`~/.sangam/models.yaml` entries; `auto · server default`
  unless you pick one); the choice rides the agent run as a per-run override
  and persists in `localStorage`.
- **Attachments** — 📎 or drag-&-drop files onto the composer; they render as
  removable chips and are delivered to Agent prompts as `context.attachments`
  (client-side read, 32 k chars/file cap — no server writes).
- **Explorer** — a lazy file tree (directories via `exec ls`, files via the
  `list_files` tool) with git-status badges (`git status --porcelain`); click a
  file to preview it (HTML renders in a sandboxed iframe).
- **Changes** — type-check diagnostics (errors/warnings with `file:line`) and
  test output from the Changes tab or the composer's Type check / Tests modes,
  **plus a unified-diff view of the working tree** (per-file `+n / −n` counts
  from `git diff --numstat`, the whole diff or one file at a time).
- **Rich replies** — assistant answers render as **markdown** (headings, lists,
  nested lists, tables, blockquotes, task lists) with **syntax-highlighted
  code blocks** (Python, JS/TS, shell, JSON, HTML/CSS, diffs). Consecutive tool
  steps collapse into one disclosure row with a counts summary. A `<think>…</think>`
  block renders as a collapsed **reasoning row** that expands to the full
  reasoning.
- **Light / dark / system theme** — a switcher in the top bar plus a 10–22px
  content font-size stepper. Both persist in `localStorage`
  (`sangam.theme`, `sangam.fontSize`) and the choice is applied by a
  synchronous inline `<head>` bootstrap, so a dark-mode user never sees a white
  flash. `system` follows the OS live.
- **App shell** — a 280px sidebar (drag-resizable 264–420px, collapsing to a
  56px icon rail automatically below 1024px, `Mod+B` to toggle) and a
  right-hand panel that opens at 45% of the window, keeps your dragged width up
  to a 70% cap, squeezes to 300px, then closes rather than squeezing the chat
  below its 400px floor.
- **Agents** — the last run's goal, plan steps with status chips, and outcome.
- **Trajectory** — a turn-aware event ledger over the sessions DB (item 6.9):
  virtualized User/Assistant/Tool rows with turn and step boundaries, a
  "Between turns" section for compaction records, an interactive timing
  **Overview** (real start/duration left→right, TTFT divided from decoding,
  hover for the exact clock, drag to focus the ledger on an interval, wheel to
  zoom, right-click to clear) and a **record inspector** (token usage,
  duration, input/output, timing, attachment summaries, and a Code view for
  tool arguments and output). It reads
  `GET /api/sessions/<id>/trajectory` and live-appends from the additive
  `record` / `token` SSE frames — see
  [session-trajectory-view](docs/features/session-trajectory-view.md).
- **⌘K palette** — FTS5 search over all chat history (`/api/search`) plus quick
  commands (new session, toggle panel, refresh changes, …).
- **Status bar** — core health/version, session + tool counts.

Everything is rendered through DOM nodes with `textContent` — `innerHTML` is
used nowhere in the webapp — so markup inside a model reply (or a diff) is
displayed as text, never executed.

Notes: since item 6.9 the agent endpoints accept an optional `session_id`, so
web turns are persisted into SQLite (the user prompt, every assistant step and
tool result, with token/latency/first-token accounting) and the Trajectory view
has something to show; the `localStorage` transcript remains as a fallback for
sessions created before that change. Server sessions (CLI/agent runs) load
read-only.
<!-- 6.9 -->
Notes: web-chat transcripts persist client-side (`localStorage`) because the
HTTP API has no append-message endpoint; server sessions (CLI/agent runs) load
read-only. `sangam web` binds to `127.0.0.1` only. `?api=<url>` (or
`localStorage.sangam.apiBase`) points the UI at a different core, but full
support requires the `sangam web` same-origin proxy (the API sends no CORS
headers).

Full task breakdown in [`todo.md`](todo.md), Phase 6.

### Testing the webapp

- `pytest tests/test_webapp_surface.py` — the HTTP surface: static serving of
  all five assets, same-origin proxy (including the SSE agent stream), the
  client contract, the `--sgw-*` token values, the app-frame geometry numbers
  and the session-description serializer.
- `pytest -m simulation tests/simulation/test_sim_webapp_browser.py` —
  **browser-driven simulations**: launches the real stack and drives
  `index.html` in a real Chromium-family browser (Edge/Chrome, found
  automatically; `$SANGAM_SIM_BROWSER` to point elsewhere) over a
  stdlib-only DevTools-Protocol client — no Playwright/Selenium. Prompts of
  different types (agent goal, exec command, type check path, tests path,
  history search) are typed into the composer and validated against the
  rendered DOM; a real-LLM tier runs the agent prompt through a local model
  endpoint when one is reachable (skips otherwise). Item 6.8 added theme
  switching, markdown/diff/reasoning rendering, the responsive ladder and an
  XSS regression to that suite (39 scenarios). Every test also fails on
  uncaught page JS errors.

## Desktop App — Roadmap

A **cross-platform native desktop app** (macOS + Windows) is a target surface for
the same harness. It keeps the existing **Python core** as the engine and adds a
**React** UI on top — the same "core is a library, surfaces are thin" split used
for the CLI and the VS Code extension:

- **Python core** — the `sangam` package, frozen into a standalone per-OS
  binary and run headless (`sangam serve`, the same HTTPS API server the VS
  Code extension talks to). All agent/tool/runtime logic stays in Python.
- **React UI** — a TypeScript + React frontend (chat panel, file tree,
  diagnostics, agent-step viewer, command palette) that speaks HTTPS to the
  core over localhost (optionally with TLS/cert configuration) — the same
  wire protocol as the VS Code extension.
- **Native shell** — a desktop wrapper that bundles the React build and spawns
  the frozen Python core as a sidecar process. Ships a native `.app`/`.dmg` on
  macOS and `.exe`/`.msi` on Windows (Tauri preferred for the OS-webview native
  feel and small binary; Electron as a fallback).

Planned packaging steps (full task breakdown in [`todo.md`](todo.md), Phase 7):

1. **React UI** — scaffold `desktop-ui/` (Vite + TypeScript + React); build the
   app shell and a typed HTTPS client to the core server.
2. **Native shell** — scaffold `desktop-app/` (Tauri or Electron); configure it
   to bundle and spawn the frozen Python core sidecar; wire frontend↔core IPC over
   localhost HTTPS.
3. **Python core packaging** — freeze the core with PyInstaller/Nuitka per OS
   (including mypy + deps); confirm the headless `serve` entry point runs the
   agent loop end-to-end.
4. **macOS build** — universal `.app` (arm64 + x86_64), `.dmg` installer,
   Developer ID code signing, Apple notarization + stapling, clean-launch test.
5. **Windows build** — `.exe` + `.msi` installer (WiX/Tauri), WebView2 runtime
   handling, code signing, clean-launch test.
6. **Release CI** — GitHub Actions matrix (macOS arm64/x86_64, Windows x64):
   build → sign → notarize (macOS) → package → attach to GitHub Releases.
7. **Updates & polish** — auto-update, native menus/tray, first-run onboarding
   (LLM creds + runtime detection), native window-state persistence.

This builds on the headless HTTPS API server from Phase 4 — the desktop app is
another thin surface over the same server, not a separate backend.

## Installation

```bash
pip install -e .
```

Requires Python >= 3.9 and `mypy>=0.900` (installed automatically).

### Headless server install (`install.sh`)

To install sangam and run it as a headless HTTP API server (the backend for the
VS Code extension, local webapp, and desktop app), use `install.sh`:

```bash
./install.sh                 # install (or update) and run the server
./install.sh --install-only  # install/update but do not start the server
./install.sh --port 8765     # bind a specific port (default 8765)
./install.sh --local         # install from THIS checkout (current branch,
                             #   working tree as-is) instead of GitHub
./install.sh --local ~/path/to/sangam   # ...from an explicit checkout
```

`install.sh` pulls the source from GitHub `main`, creates a venv at
`~/.sangam/venv`, installs the package, and runs `sangam serve`. With
`--local` it skips the clone entirely and installs **editable** (`pip -e`)
from the local checkout — the current branch and uncommitted changes — so
edits there are picked up by the installed `sangam` without reinstalling
(ideal for testing a feature branch; run `sangam web` afterwards to try it).
The server
needs a model to drive the agent, so before starting it `install.sh` **probes
for a local model runtime**:

- It queries the OpenAI-compatible `/v1/models` endpoint of **Ollama**
  (`127.0.0.1:11434`) and then **LM Studio** (`127.0.0.1:1234`).
- If a reachable endpoint advertises a model, the server is started using that
  local model.
- **sangam never installs a model runtime.** If no Ollama or LM Studio endpoint
  is reachable, `install.sh` errors out with **"no model available/selected"**
  and exits without starting the server.

Override the probe with environment variables:

```bash
SANGAM_MODEL_ENDPOINT=http://127.0.0.1:12345 SANGAM_MODEL_NAME=bonsai-27b-mlx ./install.sh
```

`--install-only` installs without requiring a model.

## Usage

### CLI

```bash
# Run a single task
sangam "read main.py"

# Interactive mode
sangam --interactive

# Batch mode (from a file)
sangam --batch tasks.txt

# Verbose output
sangam --verbose "type check src/"

# Run the coding agent on a goal
sangam agent "fix the type errors in src/ and run the tests"

# List the models sangam can use (configured + discovered)
sangam models
sangam models --no-probe          # configured models only (no network)
sangam models --json              # machine-readable listing
sangam models --select fast-local # show the client for a named entry

# Run a command in a chosen runtime
sangam exec --runtime python "print('hello')"
sangam exec --runtime node "console.log('hello')"
sangam exec --runtime shell "ls -la"

# List stored chat sessions (id, updated, surface, description/title)
sangam session list
sangam session list --limit 10 --json

# Generate one-line descriptions for sessions (via the configured model;
# falls back to the first user message when no model is available)
sangam session generate-description <session-id>
sangam session generate-description --all        # only sessions missing one
sangam session generate-description --all --force  # regenerate existing

# Continue a past session in interactive mode (prior transcript is replayed;
# new turns append to the same session)
sangam restore <session-id>

# Run one prompt through the chat pipeline (tool rounds included) and exit —
# the turn persists as a session, so `sangam session list` shows it
sangam run "summarize the failures in the last test run"
sangam run --model fast-local "what changed since the last commit?"

# Show the project instructions (AGENTS.md) loaded into the system prompt
sangam agents                      # path, size, truncation notice, preview
sangam agents --path               # just the resolved path (script-friendly)
sangam agents --preview 0          # the whole file
sangam agents --json               # machine-readable
sangam agents ./OTHER.md           # inspect an explicit file instead

# Inside `sangam -i`: free text always goes to the model; tools run directly
# only via the slash command (display-only, not added to the transcript)
#   /tool                          # list registered tools + aliases
#   /tool get_time {}              # run a tool with JSON args, no model round
#   /tool web_search {"query": "…"}
#   /history                       # legacy task-history views (moved from the
#   /search <query>                # bare `history` / `search` words in 2.40)
#   /agents                        # show the AGENTS.md loaded into the prompt
#   /agents <path>                 # switch to a different file for this session
#   /agents off | on               # drop the instructions / re-discover them
```

### Project instructions (`AGENTS.md`)

sangam loads the repository's standing instructions into the model's
system prompt, so every chat session and every `sangam agent` run follows
your project's rules instead of a generic persona alone. Discovery is
nearest-first (`AGENTS.md`, then `AGENT.md`): the search starts at the
current directory and walks up to the filesystem root, so a nested
package's own file shadows the repo root's.

Resolution order — the same for every surface (chat session, agent loop,
`sangam agents`):

1. an explicit path (`ChatSession(agents_path=…)`, `Agent(agents_path=…)`,
   `/agents <path>`, `sangam agents <path>`),
2. the `SANGAM_AGENTS_FILE` environment variable,
3. walk-up discovery from the current directory.

The content is injected with a framing line that marks it as standing
rules for this repository (without it the model treats the file as
reference material), and is capped at 20 000 characters — head and tail
are kept, the middle is replaced by a `[... truncated N chars ...]`
marker — so a huge file cannot crowd out the conversation. The session
banner names the loaded file, and `/agents` shows, switches, drops or
re-discovers it for the live session.

```bash
export SANGAM_AGENTS_FILE=~/projects/myrepo/CONVENTIONS.md   # override discovery
```

### Streamed thinking

When the backend streams reasoning, `sangam -i` shows it as one
contiguous block instead of one ragged line per token:

```
│ thinking · qwen3:0.6b the user asked about the failing login test, so I
│ should look at the auth module first and check the session store schema

Here is the answer…
```

Deltas are buffered and a line is emitted only once it is complete, then
wrapped to the terminal width — measured in rendered cells, so ANSI
escapes and double-width glyphs count correctly — with a `│ ` gutter on
every line and the model name on the first. Inline markdown in the
reasoning (`code`, **bold**, lists) renders the same way it does in the
answer, and fenced code is left verbatim. The block always ends with a
blank line, so the answer starts on its own line — including when a turn
is cancelled with Ctrl+C.

Turn it off with `/config set thinking_stream off`. Under `NO_COLOR`, or
when output is not a terminal, the block is still wrapped and guttered but
carries no colour escapes, so captured logs stay clean.

### Output channels

The three regions of a turn — tool calls, the streamed thinking block and
the final reply — each print on their own background fill, so a long turn
reads as three separable blocks instead of one mass of text:

```
[ navy   ]  > Tool call · web_search — `{"query": "urvasi song"}`
[ navy   ]  > Tool result · web_search
[ indigo ]  │ thinking · qwen3:0.6b
[ indigo ]  │ the user wants a song, so I should search and summarise
[ green  ]  I found three matches for that song:
```

The fills come from the terminal's 6-level colour cube, so a 256-colour
terminal renders them exactly and the three channels stay distinct. Colour
appears only on a colour terminal: under `NO_COLOR`, `TERM=dumb`, in CI or
when output is piped, the blocks are byte-identical to the plain text (the
gutters and labels still separate them).

```bash
/config set channel_backgrounds off      # back to the glyph-only look
/config set channel_theme light          # for a light terminal background
SANGAM_THEME=light sangam -i             # same, without touching the config
```

Display only — transcripts, `sangam session list` and `sangam restore` are
unaffected. See
[`docs/features/cli-output-channels.md`](docs/features/cli-output-channels.md).

### Model catalog (`~/.sangam/models.yaml`)

`~/.sangam/models.yaml` is the source of truth for which models sangam can use.
It follows Continue's `config.yaml` `models` shape:

```yaml
default: fast-local          # optional: the active chat/agent model

models:
  - name: fast-local
    provider: ollama
    model: qwen3:0.6b
    apiBase: http://127.0.0.1:11434/v1
    roles: [chat, agent]
    capabilities: [tool_use]
    defaultCompletionOptions:
      temperature: 0.2
      maxTokens: 2048
    requestOptions:
      timeout: 60  # seconds

  - name: remote-gpt
    provider: openai
    model: gpt-4o-mini
    apiKey: $OPENAI_API_KEY   # literal, $VAR, or ${VAR}
    roles: [chat]
```

Selection order: the file's `default` entry, else the first entry with role
`chat`, else the first entry. `sangam models` merges those configured entries
with the model ids advertised by probed OpenAI-compatible runtimes (Ollama on
`127.0.0.1:11434`, LM Studio on `127.0.0.1:1234`, and every configured
`apiBase`) via `/v1/models`. Probing is read-only and best-effort — sangam never
installs a model runtime. Override the catalog path with `SANGAM_MODELS_FILE` or
`sangam models --catalog <path>`.

`requestOptions` carries transport-level settings for the entry's HTTP calls.
`timeout` (seconds) is honoured on every chat-completions request and overrides
the `SANGAM_LLM_TIMEOUT` environment variable (which defaults the timeout to
120 s). Invalid or non-positive values are ignored.

### As a library

```python
from sangam import SangamCLI

cli = SangamCLI()
result = cli.execute_task("type check main.py")
print(result)
```

The runtime contracts are importable for building backends:

```python
from sangam.core import Executor, RunSpec, RunResult
```

The agent harness is importable too:

```python
from sangam.core.agent import Agent
from sangam.core.planner import Planner
from sangam.core.tools import ToolRegistry
from sangam.core.session_store import SessionStore
from sangam.core.model import (
    ModelClient,
    OpenAIModelClient,
    LocalModelClient,
    build_model_client,
)
```

## Theming

Every sangam surface resolves its theme through **one ordered chain**:

```
--theme flag  →  SANGAM_THEME env  →  stored preference  →  system
```

**`system` is the fallback.** With nothing stored and nothing passed you get
whatever your operating system is set to, and the webapp *follows it live* —
flip your OS to dark mode and the page repaints without a reload.

| Surface | How to set it | Remembered in |
|---|---|---|
| **CLI** | `SANGAM_THEME=light sangam -i`, or `sangam --theme light` | `~/.sangam/config.json` → `theme` (`/config set theme dark`) |
| **Webapp** | the ☀ / ☾ / ⌘ switcher, or `sangam web --theme dark` (`$SANGAM_WEB_THEME`) | the browser (`localStorage`) |
| **VS Code** | nothing to set — the panel follows the editor | — |

A few rules worth knowing:

- **`NO_COLOR` wins over everything** on the terminal — no colour at all,
  whatever the theme says.
- **The CLI's effect is deliberately small.** `dark` (and the `system`
  fallback, since a terminal can't ask the OS) is exactly how sangam has
  always looked. `light` only lifts the *dim* tier one step so secondary text
  stays legible on a light-background terminal. No palette was rewritten.
- **Nothing throws.** A typo like `SANGAM_THEME=sepie` warns once and falls
  through to the next source instead of breaking your shell session.
- **The webapp flips instantly** — a theme change is a repaint, never a
  cross-fade.

Full contract, the per-surface precedence table and how to add a palette
token: [`docs/features/theming.md`](docs/features/theming.md).

## Project Structure

```
sangam/
├── __init__.py            # Package exports
├── cli/
│   ├── main.py            # CLI surface (argparse, REPL, batch mode)
│   ├── theme.py           # the shared theme contract: resolution order,
│   │                      #   NO_COLOR + TTY detection, the dim tier
│   └── commands/
│       ├── agent.py       # sangam agent "<goal>"
│       ├── models.py      # sangam models (list/select from models.yaml)
│       ├── exec.py        # sangam exec --runtime <rt> <command>
│       ├── serve.py       # sangam serve (headless HTTP API server)
│       └── web.py         # sangam web [--theme light|dark|system]
├── core/
│   ├── engine.py          # Core engine (config + logger + task engine wiring)
│   ├── agent.py           # Agent orchestrator (plan → act → observe → reflect)
│   ├── planner.py         # Goal → step plan
│   ├── memory.py          # Conversation history + working-set
│   ├── session_store.py   # SQLite sessions/messages/usage + FTS5 search
│   ├── sandbox.py         # SandboxPolicy, SandboxedExecutor, resource limits
│   ├── model/
│   │   ├── __init__.py    # exports + build_model_client(config) factory
│   │   ├── client.py      # ModelClient protocol + data contracts
│   │   ├── catalog.py     # ~/.sangam/models.yaml catalog (load/select/probe)
│   │   ├── model.py       # OpenAIModelClient (OpenAI-compatible)
│   │   └── local.py       # LocalModelClient (local providers, e.g. Ollama/LM Studio)
│   ├── tools/
│   │   ├── registry.py    # Tool, ToolResult, ToolRegistry, JSON-schema gen
│   │   ├── file_tools.py  # read_file, write_file, edit_file, list_files, search
│   │   ├── exec_tools.py  # run_python, run_js, run_container, run_shell, install_deps
│   │   ├── type_check_tools.py # type_check, run_tests, suggest_fixes
│   │   └── web_tools.py   # web_search — web search + page fetch
│   └── runtime/
│       ├── base.py        # Executor protocol, RunSpec, RunResult contracts
│       ├── venv.py        # VenvExecutor (isolated venv + pip + run Python)
│       ├── nvm.py         # NvmExecutor (nvm/npm, JS/React)
│       ├── shell.py       # ShellExecutor (native bash/zsh, detect_shell)
│       ├── container.py   # ContainerExecutor (Docker/Podman)
│       ├── subprocess.py  # SubprocessExecutor (plain subprocess fallback)
│       ├── detect.py      # Runtime auto-detection → executor
│       └── registry.py    # EnvironmentRegistry (~/.sangam/environments.json)
├── server/                # Headless HTTP API server (Flask)
│   ├── __init__.py        # Flask app exports
│   └── http_api.py        # create_app factory, per-request SessionStore, SSE
├── task_engine.py         # Natural-language task parsing and execution
├── sangam_integration.py  # mypy subprocess wrapper and JSON parsing
├── file_ops.py            # File read/write/analysis operations
├── logger.py              # Logging and JSONL task history
├── config.py              # Configuration management
└── mypy_cli.py            # Backward-compatible shim (re-exports cli/main + core)
tests/                     # pytest test suite
docs/                      # Architecture docs
```

## Running Tests

The test suite uses [pytest](https://docs.pytest.org/). Install the test
dependencies if you don't have them:

```bash
python3 -m venv venv
./venv/bin/pip install -r requirements-dev.txt
```

Then run the tests from the project root — the wrapper runs them **in
parallel** with [pytest-xdist](https://pytest-xdist.readthedocs.io/):

```bash
./run_tests.sh                    # pytest -q -n <2 x CPU> --dist load
```

**Always go through `./run_tests.sh`** rather than calling `pytest` yourself:
the wrapper resolves the venv interpreter, checks that `pytest-xdist` is
importable, and picks the worker count. That count is **2 × the CPU count**,
not one per core — a worker in this suite spends most of its time in
`subprocess`, `urllib`, SQLite and `time.sleep` rather than on the CPU, so one
worker per core leaves cores idle waiting on I/O.

Measured on a 10-core machine, the unit subset takes **66.9 s** with 10
workers and **62.3 s** with 20 — a modest ~7%, with an identical pass/fail
set. (For scale: the same subset runs ~158 s serially, and going parallel at
all took the whole suite from ~419 s to ~99 s when the runner landed.) A run
this short spends a noticeable share of its time on worker startup and
teardown, which is why doubling the workers buys less than the theory
suggests. Runner options:

| Command | What it runs |
|---------|--------------|
| `./run_tests.sh` | whole suite, parallel (`-n <2 x CPU> --dist load`) |
| `./run_tests.sh -k session` | extra args are passed straight to pytest |
| `./run_tests.sh -n 4` | fixed worker count |
| `./run_tests.sh --dist loadfile` | group all tests from one file on one worker |
| `./run_tests.sh --serial` | one worker, no xdist (same as plain pytest) |
| `./run_tests.sh --simulation` | the live simulation suite (see below) |
| `./run_tests.sh --print-command` | print the pytest argv, run nothing |

Plain pytest is the deliberate exception, for the two cases the wrapper cannot
serve: a single process under `--pdb` (xdist disables itself there), and a
one-file `pytest <file>` poke while iterating on a single test.

`--dist` selects how work is handed out: `load` (default, any pending test to
any free worker — fastest), `loadfile` (a whole file stays together),
`loadscope` (a whole class/module stays together), `each` (runs everything on
every worker, for race testing). `SANGAM_TEST_WORKERS`, `SANGAM_TEST_DIST` and
`SANGAM_TEST_SIM_DIST` set the defaults from the environment.

Plain pytest still works, and is what to reach for when you want a single
process (`--pdb` disables xdist automatically):

```bash
python -m pytest tests/ -v
pytest tests/
pytest tests/test_file_ops.py
pytest tests/test_file_ops.py::TestFileOperations::test_read_file
```

The mypy integration tests mock the `mypy` subprocess, so they run even if mypy is not installed.

## Simulation Tests

Beyond the unit tests, `sangam` ships a **live simulation suite** that drives
real runtimes end-to-end — no mocking of venv, subprocess, or the model
endpoint. These live in `tests/simulation/` and are tagged with the
`simulation` pytest marker, so they are excluded from the default run and only
execute when requested.

**What they cover**

- `test_sim_venv_executor.py` — `VenvExecutor`: real venv creation (idempotent
  `prepare`), real pip installs of a local package, venv-vs-subprocess isolation
  proof, `run_python` convenience + env passthrough, live timeout enforcement,
  failure reporting, and sandbox-wrapped runs (allowed/disallowed cwd, proxy
  clearing).
- `test_sim_shell_executor.py` — `ShellExecutor`: real-life `.sh` fixture
  scripts (`simulate_release.sh` release pipeline, `analyze_logs.sh` log
  analysis), one-off commands (pipelines, env, stdin, cwd), multi-line
  `run_script` constructs, live timeouts, and sandbox policy.
- `test_sim_subprocess_executor.py` — `SubprocessExecutor`: real command
  execution (PATH/absolute resolution, mypy `--version`), failure reporting,
  stdin/env/cwd, live timeouts, sandbox enforcement (filesystem scope, default/
  max timeout clamping, proxy clearing), and **live POSIX resource-limit
  enforcement** (`fsize`, `cpu_time`, `nofile`).
- `test_sim_openai_endpoint.py` — a real OpenAI-compatible endpoint (the
  device's local Ollama at `http://localhost:11434/v1`, standing in for any
  provider the future `ModelClient` will target): `/v1/models`, chat
  completions (single/multi-turn, system prompt, usage accounting), SSE
  streaming, tool-calling, and error-path validation.

**Fixtures** — real-life usage-case files live in `tests/fixtures/real_life/`
(`simulate_release.sh`, `analyze_logs.sh`, `report_env.py`, sample logs, and a
small `src/` tree). The endpoint and model are overridable via
`SANGAM_OLLAMA_BASE_URL` and `SANGAM_OLLAMA_MODEL` (default `qwen3:0.6b`).

**Running the suite**

```bash
# All live simulation tests (parallel, grouped by file so each file keeps one worker)
./run_tests.sh --simulation

# A single backend
./run_tests.sh --simulation tests/simulation/test_sim_shell_executor.py

# Verbose runner with per-test traces, live logs, and a timestamped log file
./run_live_tests.sh
```

The simulation suite drives real shared resources (real venv/npm installs, a
real Ollama server, a real Chrome instance), so the runner defaults it to the
conservative `--dist loadfile` grouping; use `--serial` or `--dist load` if you
want to override that. It keeps the 2 × CPU worker default, but gains the least
from it — on a machine already busy with other work, drop to the core count
(`./run_tests.sh --simulation -n "$(sysctl -n hw.ncpu)"`) or go `--serial`.

The endpoint tests **skip** (never fail) when Ollama is not reachable or no
model is pulled, so the default suite stays green without it. A full run is
~60 tests (221 passed / 1 skipped across the whole suite; `pytest -m "not
simulation"` = 162 for CI).

## License

MIT
