Skip to content

Conversation and automatic compaction

Message structure

Conversation.messages stores OpenAI's official typed dictionaries directly (ChatCompletionMessageParam), with roles system / user / assistant / tool, without any extra wrapper:

conv = agent.conversation
conv.messages[0]           # {"role": "system", "content": "..."}
Method Description
set_system_message_content(content) replaces or inserts the system message at index 0 (where agent.system() lands)
add_user_message(content, images=None) with images, content becomes a list of text/image_url blocks; PIL images are converted to data URLs
append_user_message(extra) appends text to the last user message (this is how the structured-output schema gets injected)
add_agent_message(msg) writes an assistant message, automatically dropping an empty tool_calls: []
add_tool_result(tool_call_id, result) writes a role="tool" message whose content is result.value_str()
pop_last_message_if_user() rolls back the last user message on cancellation
pop_from_last_user_message(inclusive) the implementation behind /retry (exclusive) and /revise (inclusive)
clear() clears everything but keeps the system message, and resets the token counter, the compaction counters and the stored originals of compacted tool results
to_history() / content_to_html() / render_history_as_html() UI rendering and /render export
dumps() / dump(path) / loads() / load(path) serialisation; the on-disk key tokens_used is kept for backward compatibility

Tool calls are stored in the tool_calls array of the assistant message (including the argument strings), and each tool result is a separate role="tool" message linked by tool_call_id. Reasoning content is detected automatically through model.reasoning_field (reasoning or reasoning_content) and written back verbatim so that it survives across turns.

Two length metrics

Metric Source Purpose
total_tokens usage.total_tokens returned by the model the sole criterion for automatic compaction (reset to None after compaction so it does not trigger again immediately)
estimated_message_length() sums character counts over the content / text / reasoning / reasoning_content fields (recursing into nested dicts and lists) estimates the reclaimed fraction and decides between "reclaim again" and "escalate to a summary"

Tiered compaction

flowchart TD
  S["check before each step"] --> A["reclaim old tool results"]
  A --> B{"still over?"}
  B -->|no| E["round done"]
  B -->|yes| C["summarise"]
  C --> D{"still over?"}
  D -->|no| E
  D -->|yes| A

Tier 1: reclaim old tool results

Old role="tool" messages beyond the most recent keep_max (default 12) are replaced in place by a placeholder:

[Compacted, ID: <toolcall_id>. If this content is still needed, call extract_compacted_tool_result with this ID or re-run the tool.]

The original text stays in memory and the model can fetch it back with extract_compacted_tool_result(toolcall_id) — this is the key to "reclaim cheaply first, without losing information". A message is replaced only when the replacement really is shorter.

Tier 2: summary compaction

The cut point is chosen so the tail never starts with an orphaned tool result, and the last user message is force-kept in the tail when necessary (some servers require a user message to be present in the request). The summary is produced by a sub-agent:

Agent.inherit(agent, share_display=False, copy_toolbox=False, copy_command=False)
# disable auto_compact / enable_extensions, enable auto_confirm, copy the current conversation as context
# execute(max_iterations=1)

The summary is required as markdown sections (system_context, overview, key_facts, user_preferences, decisions, pending_tasks, open_questions, tone_context), kept within 1024 tokens in total. On success the system message is replaced with "summary + compaction note", total_tokens is cleared, and compaction_counter.summary_rounds increases by one (the tool-round counter goes back to zero).

When there is nothing to compact or the summary fails, the conversation is left untouched; the returned statuses are NOTHING_TO_CONDENSE / SUMMARIZE_FAILED, respectively.

Parameter overview

Policy parameters are fields of AutoCompactor (code level), while the switch and the threshold are configuration (user level):

Location Name Default Description
AutoCompactor toolcall_keep_max 12 how many recent tool results to keep
AutoCompactor summary_keep_max 16 number of trailing messages kept when summarising
AutoCompactor escalation_rounds 2 how many consecutive cheap-reclaim rounds before escalating to a summary; if the estimate after reclaiming is still above threshold × ESCALATION_RATIO, escalation happens in the same round
AutoCompactor max_retries 2 extra compaction rounds allowed when the limit is still exceeded after a summary; each further round halves both toolcall_keep_max and summary_keep_max
module constant ESCALATION_RATIO 0.95 the estimate must fall below 95% of the threshold to count as real progress; it lives at module level in compact, so it cannot be tuned per AutoCompactor instance
config auto_compact.enabled true on/off switch
config auto_compact.token_threshold 192000 triggering threshold

A custom strategy only has to subclass CompactorAbstract and implement auto_compact(agent), then pass Agent(compactor=MyCompactor()).

Manual triggering

Way Behaviour
/compact runs the tier-2 summarisation only (equivalent to compact_conversation)
/compact toolcall tier-1 reclamation only
agent.conversation.compact_toolcall(keep_max=12) programmatic tier-1 call (a Conversation method), returns ToolCallCompactResult(reclaimed_count, reclaimed_fraction)
xun.compact.compact_conversation(agent, keep_recent=16) programmatic tier-2 call (a module function, not exported at the top level), returns SummaryCompactResult