Xun Overview¶
Xun is a mini LLM agent framework: the core is about 4100 lines (src/xun/*.py, almost all of it hand-written), fully type-annotated, Python 3.12+ (PEP 695 syntax). It separates the three concerns of "model + tools + display layer" and, with a very small API, provides an agent execution loop that is runnable, extensible and embeddable.
The design goal in one sentence: use plain functions as tools, constrain the lifecycle with the type system, and isolate IO behind the display layer.
Key features¶
| Feature | Description | Details |
|---|---|---|
| Functions as tools | Plain Python functions (with type annotations and a docstring) are automatically turned into a JSON Schema and registered as tools — no decorator or base class needed | Tool system |
| Execution loop | Streaming model call → tool calls → results fed back, until an iteration issues no tool calls; supports pydantic structured output | Agent setup & execution |
| Four session entry points | xun terminal, xuns web service, xunc container, xunx multi-user multiplexed service |
Installation & entry points |
| Display layer abstraction | Three built-in implementations — console / web / silent — a custom one only needs two methods | Display & web service |
| Sub-agents | The two built-in tools agent_run / agent_run_parallel, with custom getters, depth limits and cascading cancellation |
Sub-agents & cancellation |
| Automatic conversation compaction | Old tool results are reclaimed first; if the limit is still exceeded, a summarisation pass runs, and that summary can itself be compacted recursively | Conversation & compaction |
| Hooks & session commands | 16 lifecycle hooks (most arguments can be modified in place); session commands such as /help, /compact, /policy |
Hooks & session commands |
| Extension system | Drop .py files into the extensions/ directory and they take effect automatically when each agent is initialised |
Extension system |
| Permission policy | Write confirmation + write allow-list + command allow-list + LLM risk assessment | Tool system |
The main design characteristics¶
1. An execution core with a minimal surface area¶
setup_agent() returns a ready-initialised agent; three lines of code are enough to run:
from xun import setup_agent
agent = setup_agent(default_tools=True)
print(agent.instruct("What is 2 + 3?").execute().unwrap())
execute() always returns a Result and never raises (KeyboardInterrupt and CancelledError excepted). Tool functions likewise only need to return a plain JSON-serialisable value or raise; the framework wraps them into a Result automatically.
2. Type-state: pushing lifecycle errors forward to static checking¶
Agent takes its lifecycle state as a generic parameter: Agent[T.Uninit] → Agent[T.Init] → Agent[T.Final]. Calling execute() without initialize(), or continuing to use an agent after finalize(), is rejected at the static-checking stage. See the Type-State pattern.
3. Execution and display fully decoupled¶
Inside the agent, all output goes through display_event() / info() / warning() / error() / get_choice(), and DisplayAbstract decides where it actually lands: the same logic can produce terminal text, a web chat event stream, or nothing at all (NullDisplay).
4. Cooperative cancellation, parent cancellation cascades to children¶
The cancellation token ChainedEvent can be attached beneath a parent agent's token; the execution loop checks it at every step, every streaming delta and before every tool call. Long-running tools get the same capability by calling ctx.agent.check_cancel() themselves.
5. Extensions are "trusted code"¶
Source files under $XUN_HOME/extensions/ are loaded in name order whenever an agent is initialised, with a fixed entry point setup_extension(ctx); an import or initialisation failure only produces a warning and never blocks startup — like a shell rc file: simple, but with the full privileges of the process.
6. Graduated compaction instead of one-shot truncation¶
When a threshold is exceeded, the cheap step runs first: reclaiming old tool results (the original text stays available through extract_compacted_tool_result). Only if the limit is still exceeded does it escalate into a single summarisation pass; if the limit is still exceeded after the summary, further compaction runs (2 extra rounds by default, keeping half as much each time), converging step by step.
Architecture at a glance¶
flowchart TB
ENTRY["Entry points<br/>xun · xuns · xunc · xunx"] --> CORE["Agent core<br/>loop + type-state + hooks"]
CFG["Config and extensions<br/>config.json · extensions"] --> CORE
CORE --> TOOLBOX["Tool box<br/>built-in tools + sub-agents"]
CORE --> CONV["Conversation<br/>history + auto compaction"]
CORE --> DISPLAY["Display layer<br/>terminal · web · silent"]
The sequence of a single execute() call:
sequenceDiagram
participant U as Caller
participant A as Agent
participant M as Model
participant T as Tools
U->>A: instruct("...")
U->>A: execute(schema=None, max_iterations=512)
loop Each iteration
A->>A: check_cancel, before_execution_step (triggers auto-compaction)
A->>M: chat.completions.create(stream=True)
M-->>A: text delta / reasoning delta / tool calls
A->>T: call_tool(name, arguments)
T-->>A: Result value
A->>A: write the assistant message and the tool result
end
A-->>U: Result[str or BaseModel, ErrorInfo]
Suggested reading paths¶
- Installation
- Choosing a session entry point
- Type
/helpin the terminal to see the commands, and/configto see the current configuration
The version these docs correspond to is shown in the footer note (injected at build time by the language configuration in mkdocs.yml).