Conversation and automatic compaction¶
Message structure¶
Conversation.messages stores OpenAI's official typed dictionaries directly (ChatCompletionMessageParam), with roles system / user / assistant / tool, without any extra wrapper:
conv = agent.conversation
conv.messages[0] # {"role": "system", "content": "..."}
| Method | Description |
|---|---|
set_system_message_content(content) |
replaces or inserts the system message at index 0 (where agent.system() lands) |
add_user_message(content, images=None) |
with images, content becomes a list of text/image_url blocks; PIL images are converted to data URLs |
append_user_message(extra) |
appends text to the last user message (this is how the structured-output schema gets injected) |
add_agent_message(msg) |
writes an assistant message, automatically dropping an empty tool_calls: [] |
add_tool_result(tool_call_id, result) |
writes a role="tool" message whose content is result.value_str() |
pop_last_message_if_user() |
rolls back the last user message on cancellation |
pop_from_last_user_message(inclusive) |
the implementation behind /retry (exclusive) and /revise (inclusive) |
clear() |
clears everything but keeps the system message, and resets the token counter, the compaction counters and the stored originals of compacted tool results |
to_history() / content_to_html() / render_history_as_html() |
UI rendering and /render export |
dumps() / dump(path) / loads() / load(path) |
serialisation; the on-disk key tokens_used is kept for backward compatibility |
Tool calls are stored in the tool_calls array of the assistant message (including the argument strings), and each tool result is a separate role="tool" message linked by tool_call_id. Reasoning content is detected automatically through model.reasoning_field (reasoning or reasoning_content) and written back verbatim so that it survives across turns.
Two length metrics¶
| Metric | Source | Purpose |
|---|---|---|
total_tokens |
usage.total_tokens returned by the model |
the sole criterion for automatic compaction (reset to None after compaction so it does not trigger again immediately) |
estimated_message_length() |
sums character counts over the content / text / reasoning / reasoning_content fields (recursing into nested dicts and lists) |
estimates the reclaimed fraction and decides between "reclaim again" and "escalate to a summary" |
Tiered compaction¶
flowchart TD
S["check before each step"] --> A["reclaim old tool results"]
A --> B{"still over?"}
B -->|no| E["round done"]
B -->|yes| C["summarise"]
C --> D{"still over?"}
D -->|no| E
D -->|yes| A
Tier 1: reclaim old tool results¶
Old role="tool" messages beyond the most recent keep_max (default 12) are replaced in place by a placeholder:
[Compacted, ID: <toolcall_id>. If this content is still needed, call extract_compacted_tool_result with this ID or re-run the tool.]
The original text stays in memory and the model can fetch it back with extract_compacted_tool_result(toolcall_id) — this is the key to "reclaim cheaply first, without losing information". A message is replaced only when the replacement really is shorter.
Tier 2: summary compaction¶
The cut point is chosen so the tail never starts with an orphaned tool result, and the last user message is force-kept in the tail when necessary (some servers require a user message to be present in the request). The summary is produced by a sub-agent:
Agent.inherit(agent, share_display=False, copy_toolbox=False, copy_command=False)
# disable auto_compact / enable_extensions, enable auto_confirm, copy the current conversation as context
# execute(max_iterations=1)
The summary is required as markdown sections (system_context, overview, key_facts, user_preferences, decisions, pending_tasks, open_questions, tone_context), kept within 1024 tokens in total. On success the system message is replaced with "summary + compaction note", total_tokens is cleared, and compaction_counter.summary_rounds increases by one (the tool-round counter goes back to zero).
When there is nothing to compact or the summary fails, the conversation is left untouched; the returned statuses are NOTHING_TO_CONDENSE / SUMMARIZE_FAILED, respectively.
Parameter overview¶
Policy parameters are fields of AutoCompactor (code level), while the switch and the threshold are configuration (user level):
| Location | Name | Default | Description |
|---|---|---|---|
AutoCompactor |
toolcall_keep_max |
12 | how many recent tool results to keep |
AutoCompactor |
summary_keep_max |
16 | number of trailing messages kept when summarising |
AutoCompactor |
escalation_rounds |
2 | how many consecutive cheap-reclaim rounds before escalating to a summary; if the estimate after reclaiming is still above threshold × ESCALATION_RATIO, escalation happens in the same round |
AutoCompactor |
max_retries |
2 | extra compaction rounds allowed when the limit is still exceeded after a summary; each further round halves both toolcall_keep_max and summary_keep_max |
| module constant | ESCALATION_RATIO |
0.95 | the estimate must fall below 95% of the threshold to count as real progress; it lives at module level in compact, so it cannot be tuned per AutoCompactor instance |
| config | auto_compact.enabled |
true |
on/off switch |
| config | auto_compact.token_threshold |
192000 | triggering threshold |
A custom strategy only has to subclass CompactorAbstract and implement auto_compact(agent), then pass Agent(compactor=MyCompactor()).
Manual triggering¶
| Way | Behaviour |
|---|---|
/compact |
runs the tier-2 summarisation only (equivalent to compact_conversation) |
/compact toolcall |
tier-1 reclamation only |
agent.conversation.compact_toolcall(keep_max=12) |
programmatic tier-1 call (a Conversation method), returns ToolCallCompactResult(reclaimed_count, reclaimed_fraction) |
xun.compact.compact_conversation(agent, keep_recent=16) |
programmatic tier-2 call (a module function, not exported at the top level), returns SummaryCompactResult |