Files
YiemAgent/learn_pi_agent/memory_management.md
T
2026-07-25 10:13:02 +07:00

625 lines
32 KiB
Markdown

# Memory & Context Management in Pi Agent
## 1. Architecture Overview
The agent manages memory at **two layers**:
```
┌─────────────────────────────────────────────────────────────────┐
│ Agent Harness │
│ ┌──────────────┐ ┌──────────────┐ ┌─────────────────────┐ │
│ │ Session │ │ Compaction │ │ Branch Navigation │ │
│ │ (JSONL tree)│ │ (summarize) │ │ (reset + summarize)│ │
│ └──────┬───────┘ └──────────────┘ └─────────────────────┘ │
└─────────┼───────────────────────────────────────────────────────┘
│ builds context
┌─────────▼───────────────────────────────────────────────────────┐
│ Agent Class │
│ ┌───────────────────────────────────────────────────────────┐ │
│ │ _state.messages: AgentMessage[] (linear transcript) │ │
│ │ _state.tools, systemPrompt, model │ │
│ └───────────────────────────────────────────────────────────┘ │
│ subscribe() → events → UI updates │
└─────────────────────────────────────────────────────────────────┘
```
**Key separation:**
- **In-memory** (`Agent`): linear transcript for the current run. Cleared on `reset()`.
- **On-disk** (`Session`): persistent tree of entries in JSONL files. Survives restarts.
- **Compaction**: replaces old on-disk history with an LLM-generated summary, controlling context window usage.
---
## 2. Data Flow: From Prompt to LLM Call
```
agent.prompt("Read README.md")
┌──────────────────────────┐
│ normalizePromptInput() │ → { role: "user", content: "..." }
└──────────┬───────────────┘
┌──────────────────────────┐
│ runWithLifecycle() │ → sets isStreaming=true
│ runAgentLoop() │
└──────────┬───────────────┘
┌─────────────────────────────────────────────────────────────┐
│ runLoop() — the main while(true) loop │
│ │
│ Inner loop: │
│ 1. Inject steering/follow-up messages │
│ 2. streamAssistantResponse() │
│ ┌─────────────────────────────────────────────┐ │
│ │ transformContext() │ │
│ │ AgentMessage[] → AgentMessage[] │ │
│ │ (prune, inject external context) │ │
│ └──────────────┬──────────────────────────────┘ │
│ ▼ │
│ ┌─────────────────────────────────────────────┐ │
│ │ convertToLlm() │ │
│ │ AgentMessage[] → Message[] │ │
│ │ (filter to user/assistant/toolResult only) │ │
│ └──────────────┬──────────────────────────────┘ │
│ ▼ │
│ ┌─────────────────────────────────────────────┐ │
│ │ streamFunction() → LLM provider │ │
│ │ { systemPrompt, messages, tools } │ │
│ └──────────────┬──────────────────────────────┘ │
│ ▼ │
│ Stream events: start → delta* → done │
│ ▼ │
│ Return AssistantMessage │
│ 3. Extract toolCall blocks │
│ 4. If toolCalls: executeToolCalls() │
│ → create toolResult messages │
│ → append to context │
│ 5. Check shouldStopAfterTurn / prepareNextTurn │
│ 6. Check steering/follow-up queues │
│ → Loop if more tool calls or queued messages │
│ │
│ Outer loop: │
│ → Check follow-up queue for messages after agent would │
│ stop │
└─────────────────────────────────────────────────────────────┘
```
---
## 3. Context Building (On-Disk → In-Memory)
The `AgentHarness` bridges on-disk session data to the in-memory agent loop.
```
session.buildContext()
┌─────────────────────────────────────────────────────────────┐
│ getBranch() — walk from leaf → root via parentId │
│ │
│ Tree structure: │
│ │
│ ┌──────┐ ┌──────┐ ┌──────────┐ ┌──────────┐ │
│ │msg 1 │───▶│msg 2 │───▶│ msg 3 │───▶│ msg 4 │ │
│ │user │ │assist│ │ toolCall │ │ user │ │
│ └──────┘ └──────┘ └──────────┘ └──────────┘ │
│ │
│ Path to root: [msg1, msg2, msg3, msg4] │
└──────────┬────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ defaultContextEntryTransform() — THE KEY STEP │
│ │
│ Finds latest "compaction" entry in path: │
│ │
│ ┌──────┐ ┌──────────┐ ┌──────┐ ┌──────┐ │
│ │msg 1 │ │compaction│ │msg 3 │ │msg 4 │ │
│ │user │ │summary X │ │msg 2 │ │assist│ │
│ └──────┘ └──────────┘ └──────┘ └──────┘ │
│ │ │ │ │ │
│ ├──────────────┤ │ │ │
│ │ SKIPPED │ │ │ │
│ │ (summarized)│ │ │ │
│ └──────────────┼───────────┘ │ │
│ ▼ ▼ │
│ Include compaction entry Include entries after │
│ + firstKeptEntryId the compaction point │
│ │
│ Result: [compaction, msg3, msg4] │
│ → compaction entry becomes a "compactionSummary" message │
└──────────┬────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ sessionEntryToContextMessages() │
│ │
│ For each entry: │
│ message → [message] │
│ compaction → [compactionSummary, ...retainedTail] │
│ branch_summary → [branchSummaryMessage] │
│ custom_message → [customMessage] │
│ other → [] (omitted from LLM context) │
│ │
│ Flat result: [compactionSummary, msg3, msg4] │
└──────────┬────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ deriveSessionContextState() │
│ │
│ Extracts from entries: │
│ thinkingLevel ← latest thinking_level_change or assistant │
│ model ← latest model_change or assistant │
│ activeTools ← latest active_tools_change │
└─────────────────────────────────────────────────────────────┘
```
---
## 4. Token Estimation
Before compaction can decide whether to trigger, it needs to know how many tokens the context uses.
```
estimateContextTokens(messages)
├── Has provider-reported usage on last assistant msg?
│ ├── YES → use actual usage + estimate tail
│ │ (accurate — avoids compounding error)
│ │
│ └── NO → estimate all messages from scratch
estimateTokens(message) [character heuristic: chars / 4]
├── role === "user"
│ content.length / 4
│ (images ≈ 4800 chars each)
├── role === "assistant"
│ sum of all content blocks:
│ text blocks → text.length
│ thinking blocks → thinking.length
│ toolCall blocks → name.length + JSON.stringify(args).length
├── role === "toolResult" / "custom"
│ content.length / 4
├── role === "bashExecution"
│ (command.length + output.length) / 4
└── role === "compactionSummary" / "branchSummary"
summary.length / 4
```
**Why `chars / 4`?** Rough heuristic: ~4 ASCII characters ≈ 1 token. Conservative estimate to avoid under-counting.
---
## 5. Compaction — The Core Memory Management
### 5.1 Trigger Condition
```
shouldCompact(contextTokens, contextWindow, settings)
→ contextTokens > contextWindow - reserveTokens
Defaults:
reserveTokens: 16384 (~16K tokens for summary prompt + output)
keepRecentTokens: 20000 (~20K tokens of recent history to keep)
Example (Claude with 200K context window):
Triggers when: contextTokens > 200000 - 16384 = 183616
```
### 5.2 Finding the Cut Point
```
findCutPoint(entries, startIndex, endIndex, keepRecentTokens)
│ Walk BACKWARD from endIndex
├── Accumulate estimated tokens per message
├── Stop when accumulated ≥ keepRecentTokens
├── Snap to nearest valid cut point
┌─────────────────────────────────────────────────────────────┐
│ Valid cut points (safe to split): │
│ - user message │
│ - assistant message │
│ - custom message │
│ - branch_summary │
│ │
│ NOT valid (tool results stay with their call): │
│ - toolResult message (skipped) │
│ │
│ Example: │
│ │
│ ┌────┐ ┌────┐ ┌────┐ ┌────┐ ┌────┐ ┌────┐ ┌────┐ │
│ │ u1 │ │ a1 │ │ tr1│ │ u2 │ │ a2 │ │ tr2│ │ u3 │ │
│ └────┘ └────┘ └────┘ └────┘ └────┘ └────┘ └────┘ │
│ ▲ ▲ ▲ │
│ │ │ │ │
│ └── kept └── cut └── discarded │
│ (~20K tokens) point history │
└─────────────────────────────────────────────────────────────┘
```
### 5.3 The Compaction Process
```
prepareCompaction(branchEntries, settings)
├── Find previous compaction (if any) → previousSummary
├── Estimate tokens of current context
├── findCutPoint() → firstKeptEntryId
├── Split into 3 groups:
│ │
│ ├── messagesToSummarize: entries BEFORE cut point
│ │ (these become the summary)
│ │
│ ├── retainedTail: entries AFTER cut point
│ │ (these stay verbatim)
│ │
│ └── turnPrefixMessages: if cut splits a turn
│ (the beginning of an interrupted turn)
└── Extract file operations from messagesToSummarize:
→ readFiles, modifiedFiles
compact(preparation, model, models)
├── If isSplitTurn:
│ ├── generateSummary(messagesToSummarize) → history summary
│ ├── generateTurnPrefixSummary(turnPrefixMessages) → turn context
│ └── Combine: history + "---" + turn prefix
├── Else (normal):
│ ├── Has previousSummary?
│ │ ├── YES → UPDATE_SUMMARIZATION_PROMPT (iterative)
│ │ └── NO → SUMMARIZATION_PROMPT (fresh)
│ └── Call LLM with conversation text + prompt
├── Append file operations:
│ "Files read: [...]\nFiles modified: [...]"
└── Return:
{
summary: "## Goal...\n## Progress...\n...",
firstKeptEntryId: "entry-uuid",
tokensBefore: 185000,
retainedTail: [msg3, msg4, ...],
details: { readFiles: [...], modifiedFiles: [...] }
}
```
### 5.4 Summary Format
```
The LLM generates a structured summary:
## Goal
- [What is the user trying to accomplish?]
## Constraints & Preferences
- [Any constraints, preferences, or requirements]
## Progress
### Done
- [x] [Completed tasks]
### In Progress
- [ ] [Current work]
### Blocked
- [Issues preventing progress]
## Key Decisions
- **[Decision]**: [Brief rationale]
## Next Steps
1. [Ordered list of what should happen next]
## Critical Context
- [Any data, examples, or references needed to continue]
Files read: [src/index.ts, package.json]
Files modified: [src/index.ts]
```
### 5.5 Iterative Compaction
Successive compact calls **update** the existing summary rather than replacing it:
```
Compaction 1 (at ~185K tokens):
Summary: "## Goal: Build a login page..."
firstKeptEntryId: "entry-003"
Compaction 2 (at ~185K tokens again):
previousSummary: "## Goal: Build a login page..."
→ UPDATE_SUMMARIZATION_PROMPT
→ PRESERVES existing information
→ ADDS new progress (move "In Progress" → "Done")
→ NEW summary: "## Goal: Build a login page... Add OAuth..."
firstKeptEntryId: "entry-003" (same boundary)
```
---
## 6. Session Storage — JSONL Format
```
Session file: .pi/sessions/--home-user--/2024-01-15T10-30-00_abc123.jsonl
Line 1 (header):
{"type":"session","version":3,"id":"abc123","timestamp":"2024-01-15T10:30:00.000Z",
"cwd":"/home/user/project","parentSession":"...","metadata":{}}
Line 2+ (entries, one per line):
{"type":"message","id":"e001","parentId":null,"timestamp":"...","message":{...}}
{"type":"message","id":"e002","parentId":"e001","timestamp":"...","message":{...}}
{"type":"compaction","id":"e003","parentId":"e002","timestamp":"...",
"summary":"## Goal: ...\n...","firstKeptEntryId":"e001",
"tokensBefore":185000}
{"type":"leaf","id":"e004","parentId":"e003","timestamp":"...",
"targetId":"e002"}
```
### Entry Types
| Type | LLM Context? | Purpose |
|------|-------------|---------|
| `message` | Yes | User, assistant, toolResult |
| `compaction` | Yes (as summary message) | Replaces compacted history |
| `branch_summary` | Yes (as summary message) | Summary of diverged branch |
| `leaf` | No | Points to current tree leaf |
| `thinking_level_change` | No | Tracking thinking level changes |
| `model_change` | No | Tracking model changes |
| `active_tools_change` | No | Tracking tool enable/disable |
| `custom` | No (unless projector configured) | App-defined data |
| `custom_message` | Yes | App-defined messages |
| `label` | No | Human-readable labels |
| `session_info` | No | Session name history |
---
## 7. Session Tree (Branching)
Sessions form a **tree**, not a linear log. This lets users "go back" and try a different approach.
```
Session tree:
┌───[e01]───┐
│ user: "a" │
└─────┬──────┘
┌──────────┐
│ assist 1 │
└─────┬────┘
┌──────────┐
│ toolCall │
└─────┬────┘
┌──────────┐
│ toolRes 1│
└─────┬────┘
┌──────────┐
│ user: "b"│ ← user goes back here
└─────┬────┘
┌─────┴─────┐
│ │
┌──────────┐ ┌──────────┐
│ user: "c" │ │ user: "d" │ ← branch point
└────┬─────┘ └────┬─────┘
│ │
┌────▼─────┐ ┌────▼─────┐
│ assist 2 │ │ assist 3 │ ← current leaf (d)
└──────────┘ └──────────┘
When user navigates to "user: b":
- Leaf moves from "d" back to "b"
- Branch summary generated for diverged work ("c" → "assist 2")
- New work branches from "b":
┌─────┐
│ user: "e" │ ← new branch
└─────┬─────┘
┌──────────┐
│ assist 4 │
└──────────┘
Context sent to LLM:
[compaction summary, user:b, user:e, assist:4]
→ The old "c"/"assist 2" branch is replaced by its summary
```
---
## 8. Pending Writes — Batching Session Persistence
To avoid writing every message individually during a run:
```
Agent loop events
handleAgentEvent(event)
├── message_end → pendingSessionWrites.push({ type: "message", message })
├── turn_end → flushPendingSessionWrites() (save_point)
├── agent_end → flushPendingSessionWrites()
flushPendingSessionWrites()
├── Iterate pendingSessionWrites[]
│ ├── message → session.appendMessage(msg)
│ ├── model_change → session.appendModelChange(provider, id)
│ ├── thinking_level_change → session.appendThinkingLevelChange(level)
│ ├── active_tools_change → session.appendActiveToolsChange(names)
│ ├── custom → session.appendCustomEntry(type, data)
│ ├── custom_message → session.appendCustomMessageEntry(...)
│ ├── label → session.appendLabel(targetId, label)
│ ├── session_info → session.appendSessionName(name)
│ └── leaf → session.getStorage().setLeafId(targetId)
└── Shift all writes → empty pending list
```
During the run, messages are accumulated in `pendingSessionWrites` and only flushed to disk at `save_point` (end of each turn) or `agent_end`.
---
## 9. Hooks & Extensibility — Context Control Points
The harness exposes hooks at every memory management boundary:
```
┌────────────────────────────────────────────────────────────────┐
│ Hooks │
├────────────────────────┬─────────────────────────────────────┤
│ Hook │ When │
├────────────────────────┼─────────────────────────────────────┤
│ before_agent_start │ Before each prompt, can add │
│ │ messages or modify system prompt │
├────────────────────────┼─────────────────────────────────────┤
│ context │ Before each LLM call, can prune/ │
│ │ modify AgentMessage[] │
├────────────────────────┼─────────────────────────────────────┤
│ before_provider_request│ Before each provider API call, can │
│ │ modify headers, retries, timeout │
├────────────────────────┼─────────────────────────────────────┤
│ before_provider_payload│ Before sending payload to provider,│
│ │ can modify the request body │
├────────────────────────┼─────────────────────────────────────┤
│ after_provider_response│ After receiving response, for │
│ │ logging/metrics │
├────────────────────────┼─────────────────────────────────────┤
│ tool_call │ Before tool execution, can │
│ │ return { block: true } │
├────────────────────────┼─────────────────────────────────────┤
│ tool_result │ After tool execution, can override │
│ │ content, details, isError, terminate│
├────────────────────────┼─────────────────────────────────────┤
│ session_before_compact│ Before compaction, can cancel or │
│ │ provide custom compaction result │
├────────────────────────┼─────────────────────────────────────┤
│ session_before_tree │ Before branch navigation, can │
│ │ cancel or provide custom summary │
├────────────────────────┼─────────────────────────────────────┤
│ prepareNextTurn │ Between turns, can replace context, │
│ │ model, or thinkingLevel │
├────────────────────────┼─────────────────────────────────────┤
│ shouldStopAfterTurn │ After a turn, if true the loop │
│ │ exits (agent_end, no more LLM calls)│
└────────────────────────┴─────────────────────────────────────┘
```
---
## 10. Complete Lifecycle: Long Session
```
Session starts empty
Turn 1: "Create a React component"
Context: [compaction summary (empty)]
Messages exchanged: ~2K tokens
└─ Session: [user1, assist1, toolCall, toolRes1, assist2]
Turn 2-10: Iterative development
Context: growing with each turn
Total context: ~50K tokens
└─ Session: [user1..assist2, user2..assist20]
Turn 15: Context approaching limit (~170K tokens)
shouldCompact() → true
└─ Compaction 1:
- Summarizes turns 1-12
- Keeps turns 13-15 verbatim
- Summary: "## Goal: React component ## Progress: built X, Y"
Turn 20: Context ~180K tokens
shouldCompact() → true
└─ Compaction 2 (iterative update):
- Updates existing summary with new progress
- "## Goal: React component ## Done: built X,Y ## New: added auth"
Turn 25: Context ~186K tokens → triggers compaction
└─ Compaction 3:
- Summary now covers ~22 turns of history
- Retained tail: last 20K tokens (~5 turns)
- Context window freed: ~186K → ~25K tokens
User navigates to Turn 8:
- Branch summary generated for Turns 9-25
- Leaf moves back to Turn 8
- Context: [compaction, branch_summary, turns 1-8]
User continues from Turn 8:
- New branch grows from Turn 8
- Old branch (9-25) replaced by branch_summary
Session ends, JSONL file persists on disk
Next session: loads from JSONL, rebuilds context
```
---
## 11. Token Budget Summary
```
Example: Claude Sonnet (200K context window)
┌─────────────────────────────────────────────────────────────┐
│ Context Window: 200,000 tokens │
├─────────────────────────────────────────────────────────────┤
│ Reserved for summary: 16,384 tokens │
├─────────────────────────────────────────────────────────────┤
│ Keep recent: 20,000 tokens │
├─────────────────────────────────────────────────────────────┤
│ Max context before compaction: 183,616 tokens │
│ (= 200000 - 16384) │
├─────────────────────────────────────────────────────────────┤
│ After compaction: ~25,000 tokens │
│ (20,000 tail + ~5,000 summary) │
│ → ~158,616 tokens freed │
└─────────────────────────────────────────────────────────────┘
```
---
## 12. Key Files Reference
| File | Responsibility |
|------|---------------|
| `agent-loop.ts` | Core loop, tool execution, streaming |
| `agent.ts` | Stateful `Agent` class, event system |
| `agent-harness.ts` | High-level harness, hooks, session management |
| `compaction/compaction.ts` | Token estimation, cut point, LLM summarization |
| `compaction/branch-summarization.ts` | Branch divergence summarization |
| `session/session.ts` | Session tree, entry appending, context building |
| `session/jsonl-storage.ts` | JSONL file read/write |
| `session/jsonl-repo.ts` | Session repo: create/open/list/delete/fork |
| `types.ts` | All type definitions |
| `messages.ts` | `convertToLlm()`, custom message helpers |
| `system-prompt.ts` | System prompt building |
| `skills.ts` | Skill management |