52 KiB
52 KiB
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ AGENT LOOP DIAGRAM │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ 1. INITIALIZATION │
│ │
│ Agent.prompt(user_input) │
│ │ │
│ ▼ │
│ normalizePrompt() ← Convert input to AgentMessage[] │
│ │ │
│ ▼ │
│ runPromptMessages() │
│ │ │
│ ▼ │
└─────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ 2. AGENT LOOP START (runAgentLoop) │
│ │
│ new_messages = copy(prompts) │
│ current_context.messages = vcat(context.messages, copy(prompts)) │
│ │ │
│ └─→ User messages are IMMEDIATELY added to context.messages │
│ (They are NOT in the steering queue!) │
│ │
│ emit(AgentStartEvent) │
│ emit(TurnStartEvent) │
│ │
│ for prompt in prompts: │
│ emit(MessageStartEvent(prompt)) │
│ emit(MessageEndEvent(prompt)) │
│ │
└─────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ 3. MAIN LOOP (runLoop - while true) │
│ │
│ pending_messages = get_steering_messages() │
│ │ │
│ └─→ Steering queue: messages from agent.steer() │
│ These are for CONTINUING conversation (NOT new user prompts) │
│ │
│ ┌───────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│ │ While has pending_messages OR has_tool_calls: │ │
│ │ │ │
│ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │
│ │ │ 4. PENDING MESSAGE HANDLING (steering messages only) │ │ │
│ │ │ │ │ │
│ │ │ pending_messages = get_steering() │ │ │
│ │ │ if !isempty(pending_messages): │ │ │
│ │ │ for msg in pending_messages: │ │ │
│ │ │ emit(MessageStartEvent(msg)) │ │ │
│ │ │ emit(MessageEndEvent(msg)) │ │ │
│ │ │ push to current_context.messages ← Steering messages go HERE │ │ │
│ │ │ push to new_messages │ │ │
│ │ │ pending_messages = [] │ │ │
│ │ │ │ │ │
│ │ │ Note: User messages from Agent.prompt() are ALREADY in context.messages │ │ │
│ │ │ (They were added in runAgentLoop via vcat(), not via this queue) │ │ │
│ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │
│ │ │ │
│ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │
│ │ │ 5. STREAM ASSISTANT RESPONSE │ │ │
│ │ │ │ │ │
│ │ │ message = streamAssistantResponse() │ │ │
│ │ │ ├─ transform_context (if configured) │ │ │
│ │ │ ├─ convert_to_llm(messages) → Message[] │ │ │
│ │ │ │ ┌───────────────────────────────────────────────────────────────────────────────────────┐ │ │ │
│ │ │ │ │ Converts AgentMessage[] to Message[] │ │ │ │
│ │ │ │ │ Filters: keeps user, assistant, toolResult │ │ │ │
│ │ │ │ └───────────────────────────────────────────────────────────────────────────────────────┘ │ │ │
│ │ │ ├─ stream_function(model, context) │ │ │
│ │ │ │ ┌───────────────────────────────────────────────────────────────────────────────────────┐ │ │ │
│ │ │ │ │ LLM Stream Events: │ │ │ │
│ │ │ │ │ • start → create partial AssistantMessage │ │ │ │
│ │ │ │ │ • text_start/delta/end → update partial message │ │ │ │
│ │ │ │ │ • thinking_start/delta/end → update partial message │ │ │ │
│ │ │ │ │ • toolcall_start/delta/end → update partial message │ │ │ │
│ │ │ │ │ • done → finalize message │ │ │ │
│ │ │ │ │ • error → handle error │ │ │ │
│ │ │ │ └───────────────────────────────────────────────────────────────────────────────────────┘ │ │ │
│ │ │ └─ push to current_context.messages & new_messages │ │ │
│ │ │ │ │ │
│ │ │ emit(MessageStartEvent(message)) │ │ │
│ │ │ emit(MessageEndEvent(message)) │ │ │
│ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │
│ │ │ │
│ │ if message.stop_reason in ("error", "aborted"): │ │
│ │ emit(TurnEndEvent) │ │
│ │ emit(AgentEndEvent) ← EXIT LOOP │ │
│ │ return │ │
│ │ │ │
│ │ tool_calls = filter(message.content, ToolCall) │ │
│ │ if !isempty(tool_calls): │ │
│ │ executeToolCalls() → ToolResultMessage[] │ │
│ │ for result in tool_results: │ │
│ │ push to current_context.messages │ │
│ │ push to new_messages │ │
│ │ emit(MessageStartEvent(result)) │ │
│ │ emit(MessageEndEvent(result)) │ │
│ │ │ │
│ │ emit(TurnEndEvent(message, tool_results)) │ │
│ │ │ │
│ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │
│ │ │ 6. PREPARE NEXT TURN │ │ │
│ │ │ │ │ │
│ │ │ next_turn_context = PrepareNextTurnContext(...) │ │ │
│ │ │ next_turn_snapshot = prepare_next_turn(config, next_turn_context) │ │ │
│ │ │ │ │ │
│ │ │ if !isnothing(next_turn_snapshot): │ │ │
│ │ │ update context, model, thinking_level │ │ │
│ │ │ │ │ │
│ │ │ if should_stop_after_turn(config, next_turn_context): │ │ │
│ │ │ emit(AgentEndEvent) ← EXIT LOOP │ │ │
│ │ │ return │ │ │
│ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │
│ │ │ │
│ │ pending_messages = get_steering_messages() ← Check for new steering messages │ │
│ │ │ │
│ └───────────────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│ │
│ follow_up_messages = get_follow_up_messages() │
│ │
│ if !isempty(follow_up_messages): │
│ pending_messages = follow_up_messages ← Continue loop for follow-ups │
│ continue │
│ │
│ break ← EXIT MAIN LOOP (no more pending messages) │
│ │
│ emit(AgentEndEvent(new_messages)) │
│ │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ 4. STEERING QUEUE MECHANISM │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ │
│ Steering messages are queued via agent.steer(message) │
│ They are ONLY processed at the START of a loop iteration │
│ AFTER the previous assistant turn completes │
│ │
│ Flow: │
│ user asks → agent responds → [user can steer here] │
│ │ │
│ └─→ pending_messages = get_steering() ← Steering messages injected here │
│ │
│ Follow-up messages are queued via agent.followUp(message) │
│ They run ONLY after agent would otherwise stop │
│ │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ COMPLETE CYCLE EXAMPLE: User asks → Agent responds → User asks 2nd → Agent responds │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ │
│ TURN #1: User asks "What is Julia?" │
│ ───────────────────────────────────────── │
│ 1. Agent.prompt("What is Julia?") │
│ normalizePrompt() → [UserMessage("What is Julia?")] │
│ runPromptMessages() │
│ │
│ 2. runAgentLoop() │
│ new_messages = [UserMessage("What is Julia?")] │
│ current_context.messages = vcat([...existing...], [UserMessage("What is Julia?")]) │
│ │ │
│ └─→ User message IMMEDIATELY added to context.messages (NOT via steering queue!) │
│ emit(AgentStartEvent), emit(TurnStartEvent) │
│ emit(MessageStart/End) for user message │
│ │
│ 3. runLoop() │
│ pending_messages = get_steering() = [] ← Steering queue is empty (no agent.steer() yet) │
│ │
│ 4. streamAssistantResponse() │
│ convert_to_llm([UserMessage]) → Message[] │
│ LLM call with [UserMessage] │
│ receive AssistantMessage: "Julia is a programming language..." │
│ push AssistantMessage to current_context.messages │
│ push AssistantMessage to new_messages │
│ emit(MessageStart/End) for assistant message │
│ │
│ 5. check stop_reason → continue (no tools, no error) │
│ │
│ 6. emit(TurnEndEvent) │
│ │
│ 7. prepare_next_turn() → nothing (default) │
│ │
│ 8. should_stop_after_turn() → false (default) │
│ │
│ 9. pending_messages = get_steering() = [] ← No steering messages │
│ │
│ 10. follow_up_messages = get_follow_up() = [] │
│ │
│ 11. break ← Exit main loop │
│ │
│ 12. emit(AgentEndEvent) │
│ │
│ ┌───────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│ │ Current context.messages: │ │
│ │ [UserMessage("What is Julia?"), AssistantMessage("Julia is...")] │ │
│ │ │ │
│ │ steering_queue: [] │ │
│ │ follow_up_queue: [] │ │
│ └───────────────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│ │
│ LLM SEES (convert_to_llm() filters): │ │
│ ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│ │ Messages passed to LLM API: │ │
│ │ [UserMessage("What is Julia?"), AssistantMessage("Julia is...")] │ │
│ └─────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│ │
│ TURN #2: User asks "How does it work?" │
│ ───────────────────────────────────────── │
│ 1. Agent.prompt("How does it work?") │
│ normalizePrompt() → [UserMessage("How does it work?")] │
│ runPromptMessages() │
│ │
│ 2. runAgentLoop() │
│ new_messages = [UserMessage("How does it work?")] │
│ current_context.messages = vcat([...previous..., UserMessage("How does it work?")]) │
│ │ │
│ └─→ User message added (context preserved from Turn #1) │
│ emit(AgentStartEvent), emit(TurnStartEvent) │
│ emit(MessageStart/End) for user message │
│ │
│ 3. runLoop() │
│ pending_messages = get_steering() = [] │
│ │
│ 4. streamAssistantResponse() │
│ convert_to_llm([UserMsg1, AssistantMsg1, UserMsg2]) → Message[] │
│ LLM call with FULL conversation history (context preserved!) │
│ receive AssistantMessage: "It works by..." │
│ push AssistantMessage to current_context.messages │
│ push AssistantMessage to new_messages │
│ │
│ 5. emit(TurnEndEvent), emit(AgentEndEvent) │
│ │
│ ┌───────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│ │ Current context.messages: │ │
│ │ [UserMsg1, AssistantMsg1, UserMsg2, AssistantMsg2] │ │
│ └───────────────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│ │
│ LLM SEES: │
│ ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│ │ Messages passed to LLM API: │ │
│ │ [UserMessage("What is Julia?"), │ │
│ │ AssistantMessage("Julia is..."), │ │
│ │ UserMessage("How does it work?"), │ │
│ │ AssistantMessage("It works by...")] │ │
│ └─────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ STEERING MESSAGES │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ │
│ What is a steering message? │
│ • A message (any AgentMessage type) injected via: `agent.steer(message)` │
│ • Goes into the steering queue, not immediately to context.messages │
│ │
│ How is it created? │
│ • User code calls: agent.steer(UserMessage("...")) │
│ • Or: agent.steer(AssistantMessage("...")) │
│ • Or any other AgentMessage subtype │
│ │
│ When is it processed? │
│ • At the START of the next loop iteration (line 194-202 in agent_loop.jl) │
│ • AFTER the previous assistant turn completes │
│ • BEFORE the next assistant response is streamed │
│ │
│ Why use steering? │
│ Use case 1: Tool execution result injection │
│ - Agent calls a tool (e.g., read_file, bash) │
│ - Tool returns result │
│ - You want to inject a follow-up question based on the result │
│ - agent.steer(UserMessage("Based on the file, what should we do next?")) │
│ │
│ Use case 2: Multi-turn conversation without user input │
│ - Agent responds to user │
│ - Before user types again, you want to inject a system message │
│ - agent.steer(BashExecutionMessage(...)) or custom message │
│ - This continues the conversation automatically │
│ │
│ Use case 3: Branch navigation recovery │
│ - User navigates between conversation branches │
│ - After switching branches, you want to inject a context message │
│ - agent.steer(BranchSummaryMessage(...)) │
│ - The agent can then continue from the new branch context │
│ │
│ Use case 4: Compaction summary injection │
│ - Conversation history is compacted │
│ - After compaction, inject summary message │
│ - agent.steer(CompactionSummaryMessage(...)) │
│ - Agent knows old history was summarized │
│ │
│ Example: │
│ agent.steer(UserMessage("Follow-up question here")) │
│ # This will be processed in the next loop iteration, │
│ # appearing in context.messages before the next LLM call │
│ │
│ The LLM sees: │
│ ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│ │ All messages become Message[] via convert_to_llm(): │ │
│ │ [UserMessage(...), AssistantMessage(...), UserMessage(from_steer), ...] │ │
│ │ │ │
│ │ The LLM cannot tell which came from Agent.prompt() vs agent.steer() │ │
│ └─────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│ │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ LLM PROCESSING: How LLM sees messages │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ │
│ The LLM NEVER sees "user message" vs "steering message" - it only sees Message types: │
│ │
│ ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│ │ convert_to_llm() transforms ALL AgentMessages to Message[]: │ │
│ │ │ │
│ │ UserMessage("user") → UserMessage (for LLM) │ │
│ │ Steering UserMessage("user") → UserMessage (for LLM) ← Same! │ │
│ │ AssistantMessage("assistant") → AssistantMessage (for LLM) │ │
│ │ ToolResultMessage("toolResult") → ToolResultMessage (for LLM) │ │
│ │ │ │
│ │ BranchSummaryMessage → UserMessage (wrapped in summary tags) │ │
│ │ CompactionSummaryMessage → UserMessage (wrapped in summary tags) │ │
│ │ BashExecutionMessage → UserMessage (if not excluded) │ │
│ │ CustomMessage → UserMessage │ │
│ └─────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│ │
│ The difference is ONLY in HOW messages enter the system: │
│ • User messages: Agent.prompt() → vcat() → context.messages (direct) │
│ • Steering: agent.steer() → queue → loop → context.messages (indirect) │
│ │
│ At LLM level: BOTH become UserMessage in the conversation! │
│ │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ KEY INSIGHTS │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ │
│ 1. User prompts go DIRECTLY to context.messages via vcat() in runAgentLoop() │
│ │
│ 2. Steering queue is for messages injected via agent.steer() AFTER a turn finishes │
│ This allows continuing conversation without calling Agent.prompt() again │
│ │
│ 3. Context is preserved across turns - context.messages grows with each turn │
│ LLM sees the full conversation history │
│ │
│ 4. At LLM level, ALL messages become Message types (UserMessage/AssistantMessage/ToolResultMessage) │
│ The "steering" vs "user" distinction is just a control mechanism, not a message type │
│ │
│ 5. New turn is triggered by: │
│ - New Agent.prompt() call (adds user messages) │
│ - Steering messages (adds steering messages) │
│ - Follow-up messages (adds follow-up messages) │
│ │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘