Skip to content

Xun Overview

Xun is a mini LLM agent framework: the core is about 4100 lines (src/xun/*.py, almost all of it hand-written), fully type-annotated, Python 3.12+ (PEP 695 syntax). It separates the three concerns of "model + tools + display layer" and, with a very small API, provides an agent execution loop that is runnable, extensible and embeddable.

The design goal in one sentence: use plain functions as tools, constrain the lifecycle with the type system, and isolate IO behind the display layer.

Key features

Feature Description Details
Functions as tools Plain Python functions (with type annotations and a docstring) are automatically turned into a JSON Schema and registered as tools — no decorator or base class needed Tool system
Execution loop Streaming model call → tool calls → results fed back, until an iteration issues no tool calls; supports pydantic structured output Agent setup & execution
Four session entry points xun terminal, xuns web service, xunc container, xunx multi-user multiplexed service Installation & entry points
Display layer abstraction Three built-in implementations — console / web / silent — a custom one only needs two methods Display & web service
Sub-agents The two built-in tools agent_run / agent_run_parallel, with custom getters, depth limits and cascading cancellation Sub-agents & cancellation
Automatic conversation compaction Old tool results are reclaimed first; if the limit is still exceeded, a summarisation pass runs, and that summary can itself be compacted recursively Conversation & compaction
Hooks & session commands 16 lifecycle hooks (most arguments can be modified in place); session commands such as /help, /compact, /policy Hooks & session commands
Extension system Drop .py files into the extensions/ directory and they take effect automatically when each agent is initialised Extension system
Permission policy Write confirmation + write allow-list + command allow-list + LLM risk assessment Tool system

The main design characteristics

1. An execution core with a minimal surface area

setup_agent() returns a ready-initialised agent; three lines of code are enough to run:

from xun import setup_agent

agent = setup_agent(default_tools=True)
print(agent.instruct("What is 2 + 3?").execute().unwrap())

execute() always returns a Result and never raises (KeyboardInterrupt and CancelledError excepted). Tool functions likewise only need to return a plain JSON-serialisable value or raise; the framework wraps them into a Result automatically.

2. Type-state: pushing lifecycle errors forward to static checking

Agent takes its lifecycle state as a generic parameter: Agent[T.Uninit] → Agent[T.Init] → Agent[T.Final]. Calling execute() without initialize(), or continuing to use an agent after finalize(), is rejected at the static-checking stage. See the Type-State pattern.

3. Execution and display fully decoupled

Inside the agent, all output goes through display_event() / info() / warning() / error() / get_choice(), and DisplayAbstract decides where it actually lands: the same logic can produce terminal text, a web chat event stream, or nothing at all (NullDisplay).

4. Cooperative cancellation, parent cancellation cascades to children

The cancellation token ChainedEvent can be attached beneath a parent agent's token; the execution loop checks it at every step, every streaming delta and before every tool call. Long-running tools get the same capability by calling ctx.agent.check_cancel() themselves.

5. Extensions are "trusted code"

Source files under $XUN_HOME/extensions/ are loaded in name order whenever an agent is initialised, with a fixed entry point setup_extension(ctx); an import or initialisation failure only produces a warning and never blocks startup — like a shell rc file: simple, but with the full privileges of the process.

6. Graduated compaction instead of one-shot truncation

When a threshold is exceeded, the cheap step runs first: reclaiming old tool results (the original text stays available through extract_compacted_tool_result). Only if the limit is still exceeded does it escalate into a single summarisation pass; if the limit is still exceeded after the summary, further compaction runs (2 extra rounds by default, keeping half as much each time), converging step by step.

Architecture at a glance

flowchart TB
  ENTRY["Entry points<br/>xun · xuns · xunc · xunx"] --> CORE["Agent core<br/>loop + type-state + hooks"]
  CFG["Config and extensions<br/>config.json · extensions"] --> CORE
  CORE --> TOOLBOX["Tool box<br/>built-in tools + sub-agents"]
  CORE --> CONV["Conversation<br/>history + auto compaction"]
  CORE --> DISPLAY["Display layer<br/>terminal · web · silent"]

The sequence of a single execute() call:

sequenceDiagram
  participant U as Caller
  participant A as Agent
  participant M as Model
  participant T as Tools
  U->>A: instruct("...")
  U->>A: execute(schema=None, max_iterations=512)
  loop Each iteration
    A->>A: check_cancel, before_execution_step (triggers auto-compaction)
    A->>M: chat.completions.create(stream=True)
    M-->>A: text delta / reasoning delta / tool calls
    A->>T: call_tool(name, arguments)
    T-->>A: Result value
    A->>A: write the assistant message and the tool result
  end
  A-->>U: Result[str or BaseModel, ErrorInfo]

Suggested reading paths


The version these docs correspond to is shown in the footer note (injected at build time by the language configuration in mkdocs.yml).