``` ┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ AGENT LOOP DIAGRAM │ └─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ 1. INITIALIZATION │ │ │ │ Agent.prompt(user_input) │ │ │ │ │ ▼ │ │ normalizePrompt() ← Convert input to AgentMessage[] │ │ │ │ │ ▼ │ │ runPromptMessages() │ │ │ │ │ ▼ │ └─────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ 2. AGENT LOOP START (runAgentLoop) │ │ │ │ new_messages = copy(prompts) │ │ current_context.messages = vcat(context.messages, copy(prompts)) │ │ │ │ │ └─→ User messages are IMMEDIATELY added to context.messages │ │ (They are NOT in the steering queue!) │ │ │ │ emit(AgentStartEvent) │ │ emit(TurnStartEvent) │ │ │ │ for prompt in prompts: │ │ emit(MessageStartEvent(prompt)) │ │ emit(MessageEndEvent(prompt)) │ │ │ └─────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ ▼ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ 3. MAIN LOOP (runLoop - while true) │ │ │ │ pending_messages = get_steering_messages() │ │ │ │ │ └─→ Steering queue: messages from agent.steer() │ │ These are for CONTINUING conversation (NOT new user prompts) │ │ │ │ ┌───────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ While has pending_messages OR has_tool_calls: │ │ │ │ │ │ │ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ │ │ 4. PENDING MESSAGE HANDLING (steering messages only) │ │ │ │ │ │ │ │ │ │ │ │ pending_messages = get_steering() │ │ │ │ │ │ if !isempty(pending_messages): │ │ │ │ │ │ for msg in pending_messages: │ │ │ │ │ │ emit(MessageStartEvent(msg)) │ │ │ │ │ │ emit(MessageEndEvent(msg)) │ │ │ │ │ │ push to current_context.messages ← Steering messages go HERE │ │ │ │ │ │ push to new_messages │ │ │ │ │ │ pending_messages = [] │ │ │ │ │ │ │ │ │ │ │ │ Note: User messages from Agent.prompt() are ALREADY in context.messages │ │ │ │ │ │ (They were added in runAgentLoop via vcat(), not via this queue) │ │ │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ │ │ │ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ │ │ 5. STREAM ASSISTANT RESPONSE │ │ │ │ │ │ │ │ │ │ │ │ message = streamAssistantResponse() │ │ │ │ │ │ ├─ transform_context (if configured) │ │ │ │ │ │ ├─ convert_to_llm(messages) → Message[] │ │ │ │ │ │ │ ┌───────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ │ │ │ │ │ Converts AgentMessage[] to Message[] │ │ │ │ │ │ │ │ │ Filters: keeps user, assistant, toolResult │ │ │ │ │ │ │ │ └───────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ │ │ ├─ stream_function(model, context) │ │ │ │ │ │ │ ┌───────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ │ │ │ │ │ LLM Stream Events: │ │ │ │ │ │ │ │ │ • start → create partial AssistantMessage │ │ │ │ │ │ │ │ │ • text_start/delta/end → update partial message │ │ │ │ │ │ │ │ │ • thinking_start/delta/end → update partial message │ │ │ │ │ │ │ │ │ • toolcall_start/delta/end → update partial message │ │ │ │ │ │ │ │ │ • done → finalize message │ │ │ │ │ │ │ │ │ • error → handle error │ │ │ │ │ │ │ │ └───────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ │ │ └─ push to current_context.messages & new_messages │ │ │ │ │ │ │ │ │ │ │ │ emit(MessageStartEvent(message)) │ │ │ │ │ │ emit(MessageEndEvent(message)) │ │ │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ │ │ │ │ if message.stop_reason in ("error", "aborted"): │ │ │ │ emit(TurnEndEvent) │ │ │ │ emit(AgentEndEvent) ← EXIT LOOP │ │ │ │ return │ │ │ │ │ │ │ │ tool_calls = filter(message.content, ToolCall) │ │ │ │ if !isempty(tool_calls): │ │ │ │ executeToolCalls() → ToolResultMessage[] │ │ │ │ for result in tool_results: │ │ │ │ push to current_context.messages │ │ │ │ push to new_messages │ │ │ │ emit(MessageStartEvent(result)) │ │ │ │ emit(MessageEndEvent(result)) │ │ │ │ │ │ │ │ emit(TurnEndEvent(message, tool_results)) │ │ │ │ │ │ │ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ │ │ 6. PREPARE NEXT TURN │ │ │ │ │ │ │ │ │ │ │ │ next_turn_context = PrepareNextTurnContext(...) │ │ │ │ │ │ next_turn_snapshot = prepare_next_turn(config, next_turn_context) │ │ │ │ │ │ │ │ │ │ │ │ if !isnothing(next_turn_snapshot): │ │ │ │ │ │ update context, model, thinking_level │ │ │ │ │ │ │ │ │ │ │ │ if should_stop_after_turn(config, next_turn_context): │ │ │ │ │ │ emit(AgentEndEvent) ← EXIT LOOP │ │ │ │ │ │ return │ │ │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ │ │ │ │ pending_messages = get_steering_messages() ← Check for new steering messages │ │ │ │ │ │ │ └───────────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ follow_up_messages = get_follow_up_messages() │ │ │ │ if !isempty(follow_up_messages): │ │ pending_messages = follow_up_messages ← Continue loop for follow-ups │ │ continue │ │ │ │ break ← EXIT MAIN LOOP (no more pending messages) │ │ │ │ emit(AgentEndEvent(new_messages)) │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ 4. STEERING QUEUE MECHANISM │ ├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤ │ │ │ Steering messages are queued via agent.steer(message) │ │ They are ONLY processed at the START of a loop iteration │ │ AFTER the previous assistant turn completes │ │ │ │ Flow: │ │ user asks → agent responds → [user can steer here] │ │ │ │ │ └─→ pending_messages = get_steering() ← Steering messages injected here │ │ │ │ Follow-up messages are queued via agent.followUp(message) │ │ They run ONLY after agent would otherwise stop │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ COMPLETE CYCLE EXAMPLE: User asks → Agent responds → User asks 2nd → Agent responds │ ├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤ │ │ │ TURN #1: User asks "What is Julia?" │ │ ───────────────────────────────────────── │ │ 1. Agent.prompt("What is Julia?") │ │ normalizePrompt() → [UserMessage("What is Julia?")] │ │ runPromptMessages() │ │ │ │ 2. runAgentLoop() │ │ new_messages = [UserMessage("What is Julia?")] │ │ current_context.messages = vcat([...existing...], [UserMessage("What is Julia?")]) │ │ │ │ │ └─→ User message IMMEDIATELY added to context.messages (NOT via steering queue!) │ │ emit(AgentStartEvent), emit(TurnStartEvent) │ │ emit(MessageStart/End) for user message │ │ │ │ 3. runLoop() │ │ pending_messages = get_steering() = [] ← Steering queue is empty (no agent.steer() yet) │ │ │ │ 4. streamAssistantResponse() │ │ convert_to_llm([UserMessage]) → Message[] │ │ LLM call with [UserMessage] │ │ receive AssistantMessage: "Julia is a programming language..." │ │ push AssistantMessage to current_context.messages │ │ push AssistantMessage to new_messages │ │ emit(MessageStart/End) for assistant message │ │ │ │ 5. check stop_reason → continue (no tools, no error) │ │ │ │ 6. emit(TurnEndEvent) │ │ │ │ 7. prepare_next_turn() → nothing (default) │ │ │ │ 8. should_stop_after_turn() → false (default) │ │ │ │ 9. pending_messages = get_steering() = [] ← No steering messages │ │ │ │ 10. follow_up_messages = get_follow_up() = [] │ │ │ │ 11. break ← Exit main loop │ │ │ │ 12. emit(AgentEndEvent) │ │ │ │ ┌───────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ Current context.messages: │ │ │ │ [UserMessage("What is Julia?"), AssistantMessage("Julia is...")] │ │ │ │ │ │ │ │ steering_queue: [] │ │ │ │ follow_up_queue: [] │ │ │ └───────────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ LLM SEES (convert_to_llm() filters): │ │ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ Messages passed to LLM API: │ │ │ │ [UserMessage("What is Julia?"), AssistantMessage("Julia is...")] │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ TURN #2: User asks "How does it work?" │ │ ───────────────────────────────────────── │ │ 1. Agent.prompt("How does it work?") │ │ normalizePrompt() → [UserMessage("How does it work?")] │ │ runPromptMessages() │ │ │ │ 2. runAgentLoop() │ │ new_messages = [UserMessage("How does it work?")] │ │ current_context.messages = vcat([...previous..., UserMessage("How does it work?")]) │ │ │ │ │ └─→ User message added (context preserved from Turn #1) │ │ emit(AgentStartEvent), emit(TurnStartEvent) │ │ emit(MessageStart/End) for user message │ │ │ │ 3. runLoop() │ │ pending_messages = get_steering() = [] │ │ │ │ 4. streamAssistantResponse() │ │ convert_to_llm([UserMsg1, AssistantMsg1, UserMsg2]) → Message[] │ │ LLM call with FULL conversation history (context preserved!) │ │ receive AssistantMessage: "It works by..." │ │ push AssistantMessage to current_context.messages │ │ push AssistantMessage to new_messages │ │ │ │ 5. emit(TurnEndEvent), emit(AgentEndEvent) │ │ │ │ ┌───────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ Current context.messages: │ │ │ │ [UserMsg1, AssistantMsg1, UserMsg2, AssistantMsg2] │ │ │ └───────────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ LLM SEES: │ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ Messages passed to LLM API: │ │ │ │ [UserMessage("What is Julia?"), │ │ │ │ AssistantMessage("Julia is..."), │ │ │ │ UserMessage("How does it work?"), │ │ │ │ AssistantMessage("It works by...")] │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ STEERING MESSAGES │ ├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤ │ │ │ What is a steering message? │ │ • A message (any AgentMessage type) injected via: `agent.steer(message)` │ │ • Goes into the steering queue, not immediately to context.messages │ │ │ │ How is it created? │ │ • User code calls: agent.steer(UserMessage("...")) │ │ • Or: agent.steer(AssistantMessage("...")) │ │ • Or any other AgentMessage subtype │ │ │ │ When is it processed? │ │ • At the START of the next loop iteration (line 194-202 in agent_loop.jl) │ │ • AFTER the previous assistant turn completes │ │ • BEFORE the next assistant response is streamed │ │ │ │ Why use steering? │ │ Use case 1: Tool execution result injection │ │ - Agent calls a tool (e.g., read_file, bash) │ │ - Tool returns result │ │ - You want to inject a follow-up question based on the result │ │ - agent.steer(UserMessage("Based on the file, what should we do next?")) │ │ │ │ Use case 2: Multi-turn conversation without user input │ │ - Agent responds to user │ │ - Before user types again, you want to inject a system message │ │ - agent.steer(BashExecutionMessage(...)) or custom message │ │ - This continues the conversation automatically │ │ │ │ Use case 3: Branch navigation recovery │ │ - User navigates between conversation branches │ │ - After switching branches, you want to inject a context message │ │ - agent.steer(BranchSummaryMessage(...)) │ │ - The agent can then continue from the new branch context │ │ │ │ Use case 4: Compaction summary injection │ │ - Conversation history is compacted │ │ - After compaction, inject summary message │ │ - agent.steer(CompactionSummaryMessage(...)) │ │ - Agent knows old history was summarized │ │ │ │ Example: │ │ agent.steer(UserMessage("Follow-up question here")) │ │ # This will be processed in the next loop iteration, │ │ # appearing in context.messages before the next LLM call │ │ │ │ The LLM sees: │ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ All messages become Message[] via convert_to_llm(): │ │ │ │ [UserMessage(...), AssistantMessage(...), UserMessage(from_steer), ...] │ │ │ │ │ │ │ │ The LLM cannot tell which came from Agent.prompt() vs agent.steer() │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ LLM PROCESSING: How LLM sees messages │ ├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤ │ │ │ The LLM NEVER sees "user message" vs "steering message" - it only sees Message types: │ │ │ │ ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │ │ convert_to_llm() transforms ALL AgentMessages to Message[]: │ │ │ │ │ │ │ │ UserMessage("user") → UserMessage (for LLM) │ │ │ │ Steering UserMessage("user") → UserMessage (for LLM) ← Same! │ │ │ │ AssistantMessage("assistant") → AssistantMessage (for LLM) │ │ │ │ ToolResultMessage("toolResult") → ToolResultMessage (for LLM) │ │ │ │ │ │ │ │ BranchSummaryMessage → UserMessage (wrapped in summary tags) │ │ │ │ CompactionSummaryMessage → UserMessage (wrapped in summary tags) │ │ │ │ BashExecutionMessage → UserMessage (if not excluded) │ │ │ │ CustomMessage → UserMessage │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │ │ │ The difference is ONLY in HOW messages enter the system: │ │ • User messages: Agent.prompt() → vcat() → context.messages (direct) │ │ • Steering: agent.steer() → queue → loop → context.messages (indirect) │ │ │ │ At LLM level: BOTH become UserMessage in the conversation! │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ KEY INSIGHTS │ ├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤ │ │ │ 1. User prompts go DIRECTLY to context.messages via vcat() in runAgentLoop() │ │ │ │ 2. Steering queue is for messages injected via agent.steer() AFTER a turn finishes │ │ This allows continuing conversation without calling Agent.prompt() again │ │ │ │ 3. Context is preserved across turns - context.messages grows with each turn │ │ LLM sees the full conversation history │ │ │ │ 4. At LLM level, ALL messages become Message types (UserMessage/AssistantMessage/ToolResultMessage) │ │ The "steering" vs "user" distinction is just a control mechanism, not a message type │ │ │ │ 5. New turn is triggered by: │ │ - New Agent.prompt() call (adds user messages) │ │ - Steering messages (adds steering messages) │ │ - Follow-up messages (adds follow-up messages) │ │ │ └─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘ ```