Files
YiemAgent/docs/agent_loop_diagram.md
T
2026-07-29 11:14:04 +07:00

52 KiB

┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│                                              AGENT LOOP DIAGRAM                                                 │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ 1. INITIALIZATION                                                                                               │
│                                                                                                                 │
│   Agent.prompt(user_input)                                                                                      │
│         │                                                                                                       │
│         ▼                                                                                                       │
│   normalizePrompt() ← Convert input to AgentMessage[]                                                           │
│         │                                                                                                       │
│         ▼                                                                                                       │
│   runPromptMessages()                                                                                           │
│         │                                                                                                       │
│         ▼                                                                                                       │
└─────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┘
          │
          ▼
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ 2. AGENT LOOP START (runAgentLoop)                                                                              │
│                                                                                                                 │
│   new_messages = copy(prompts)                                                                                  │
│   current_context.messages = vcat(context.messages, copy(prompts))                                             │
│         │                                                                                                       │
│         └─→ User messages are IMMEDIATELY added to context.messages                                            │
│             (They are NOT in the steering queue!)                                                               │
│                                                                                                                 │
│   emit(AgentStartEvent)                                                                                         │
│   emit(TurnStartEvent)                                                                                          │
│                                                                                                                 │
│   for prompt in prompts:                                                                                        │
│     emit(MessageStartEvent(prompt))                                                                             │
│     emit(MessageEndEvent(prompt))                                                                               │
│                                                                                                                 │
└─────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┘
          │
          ▼
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ 3. MAIN LOOP (runLoop - while true)                                                                             │
│                                                                                                                 │
│   pending_messages = get_steering_messages()                                                                    │
│         │                                                                                                       │
│         └─→ Steering queue: messages from agent.steer()                                                        │
│             These are for CONTINUING conversation (NOT new user prompts)                                        │
│                                                                                                                 │
│   ┌───────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│   │ While has pending_messages OR has_tool_calls:                                                             │ │
│   │                                                                                                           │ │
│   │   ┌─────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │
│   │   │ 4. PENDING MESSAGE HANDLING (steering messages only)                                                │ │ │
│   │   │                                                                                                     │ │ │
│   │   │   pending_messages = get_steering()                                                                  │ │ │
│   │   │   if !isempty(pending_messages):                                                                    │ │ │
│   │   │     for msg in pending_messages:                                                                    │ │ │
│   │   │       emit(MessageStartEvent(msg))                                                                  │ │ │
│   │   │       emit(MessageEndEvent(msg))                                                                    │ │ │
│   │   │       push to current_context.messages ← Steering messages go HERE                                  │ │ │
│   │   │       push to new_messages                                                                         │ │ │
│   │   │     pending_messages = []                                                                           │ │ │
│   │   │                                                                                                     │ │ │
│   │   │   Note: User messages from Agent.prompt() are ALREADY in context.messages                             │ │ │
│   │   │   (They were added in runAgentLoop via vcat(), not via this queue)                                   │ │ │
│   │   └─────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │
│   │                                                                                                           │ │
│   │   ┌─────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │
│   │   │ 5. STREAM ASSISTANT RESPONSE                                                                        │ │ │
│   │   │                                                                                                     │ │ │
│   │   │   message = streamAssistantResponse()                                                               │ │ │
│   │   │     ├─ transform_context (if configured)                                                            │ │ │
│   │   │     ├─ convert_to_llm(messages) → Message[]                                                         │ │ │
│   │   │     │   ┌───────────────────────────────────────────────────────────────────────────────────────┐   │ │ │
│   │   │     │   │ Converts AgentMessage[] to Message[]                                                  │   │ │ │
│   │   │     │   │ Filters: keeps user, assistant, toolResult                                            │   │ │ │
│   │   │     │   └───────────────────────────────────────────────────────────────────────────────────────┘   │ │ │
│   │   │     ├─ stream_function(model, context)                                                              │ │ │
│   │   │     │   ┌───────────────────────────────────────────────────────────────────────────────────────┐   │ │ │
│   │   │     │   │ LLM Stream Events:                                                                    │   │ │ │
│   │   │     │   │   • start → create partial AssistantMessage                                           │   │ │ │
│   │   │     │   │   • text_start/delta/end → update partial message                                     │   │ │ │
│   │   │     │   │   • thinking_start/delta/end → update partial message                                 │   │ │ │
│   │   │     │   │   • toolcall_start/delta/end → update partial message                                 │   │ │ │
│   │   │     │   │   • done → finalize message                                                           │   │ │ │
│   │   │     │   │   • error → handle error                                                              │   │ │ │
│   │   │     │   └───────────────────────────────────────────────────────────────────────────────────────┘   │ │ │
│   │   │     └─ push to current_context.messages & new_messages                                              │ │ │
│   │   │                                                                                                     │ │ │
│   │   │   emit(MessageStartEvent(message))                                                                  │ │ │
│   │   │   emit(MessageEndEvent(message))                                                                    │ │ │
│   │   └─────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │
│   │                                                                                                           │ │
│   │   if message.stop_reason in ("error", "aborted"):                                                         │ │
│   │     emit(TurnEndEvent)                                                                                    │ │
│   │     emit(AgentEndEvent) ← EXIT LOOP                                                                       │ │
│   │     return                                                                                                │ │
│   │                                                                                                           │ │
│   │   tool_calls = filter(message.content, ToolCall)                                                          │ │
│   │   if !isempty(tool_calls):                                                                                │ │
│   │     executeToolCalls() → ToolResultMessage[]                                                              │ │
│   │     for result in tool_results:                                                                           │ │
│   │       push to current_context.messages                                                                    │ │
│   │       push to new_messages                                                                                │ │
│   │       emit(MessageStartEvent(result))                                                                     │ │
│   │       emit(MessageEndEvent(result))                                                                       │ │
│   │                                                                                                           │ │
│   │   emit(TurnEndEvent(message, tool_results))                                                               │ │
│   │                                                                                                           │ │
│   │   ┌─────────────────────────────────────────────────────────────────────────────────────────────────────┐ │ │
│   │   │ 6. PREPARE NEXT TURN                                                                                │ │ │
│   │   │                                                                                                     │ │ │
│   │   │   next_turn_context = PrepareNextTurnContext(...)                                                   │ │ │
│   │   │   next_turn_snapshot = prepare_next_turn(config, next_turn_context)                                 │ │ │
│   │   │                                                                                                     │ │ │
│   │   │   if !isnothing(next_turn_snapshot):                                                                │ │ │
│   │   │     update context, model, thinking_level                                                           │ │ │
│   │   │                                                                                                     │ │ │
│   │   │   if should_stop_after_turn(config, next_turn_context):                                             │ │ │
│   │   │     emit(AgentEndEvent) ← EXIT LOOP                                                                 │ │ │
│   │   │     return                                                                                          │ │ │
│   │   └─────────────────────────────────────────────────────────────────────────────────────────────────────┘ │ │
│   │                                                                                                           │ │
│   │   pending_messages = get_steering_messages() ← Check for new steering messages                            │ │
│   │                                                                                                           │ │
│   └───────────────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│                                                                                                                 │
│   follow_up_messages = get_follow_up_messages()                                                                 │
│                                                                                                                 │
│   if !isempty(follow_up_messages):                                                                              │
│     pending_messages = follow_up_messages ← Continue loop for follow-ups                                        │
│     continue                                                                                                    │
│                                                                                                                 │
│   break  ← EXIT MAIN LOOP (no more pending messages)                                                            │
│                                                                                                                 │
│   emit(AgentEndEvent(new_messages))                                                                             │
│                                                                                                                 │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ 4. STEERING QUEUE MECHANISM                                                                                     │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│                                                                                                                 │
│   Steering messages are queued via agent.steer(message)                                                         │
│   They are ONLY processed at the START of a loop iteration                                                      │
│   AFTER the previous assistant turn completes                                                                   │
│                                                                                                                 │
│   Flow:                                                                                                         │
│     user asks → agent responds → [user can steer here]                                                          │
│                     │                                                                                           │
│                     └─→ pending_messages = get_steering() ← Steering messages injected here                     │
│                                                                                                                 │
│   Follow-up messages are queued via agent.followUp(message)                                                     │
│   They run ONLY after agent would otherwise stop                                                                │
│                                                                                                                 │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ COMPLETE CYCLE EXAMPLE: User asks → Agent responds → User asks 2nd → Agent responds                             │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│                                                                                                                 │
│   TURN #1: User asks "What is Julia?"                                                                           │
│   ─────────────────────────────────────────                                                                     │
│   1. Agent.prompt("What is Julia?")                                                                             │
│      normalizePrompt() → [UserMessage("What is Julia?")]                                                        │
│      runPromptMessages()                                                                                        │
│                                                                                                                 │
│   2. runAgentLoop()                                                                                             │
│      new_messages = [UserMessage("What is Julia?")]                                                             │
│      current_context.messages = vcat([...existing...], [UserMessage("What is Julia?")])                         │
│         │                                                                                                       │
│         └─→ User message IMMEDIATELY added to context.messages (NOT via steering queue!)                       │
│      emit(AgentStartEvent), emit(TurnStartEvent)                                                                │
│      emit(MessageStart/End) for user message                                                                    │
│                                                                                                                 │
│   3. runLoop()                                                                                                  │
│      pending_messages = get_steering() = []  ← Steering queue is empty (no agent.steer() yet)                   │
│                                                                                                                 │
│   4. streamAssistantResponse()                                                                                  │
│      convert_to_llm([UserMessage]) → Message[]                                                                  │
│      LLM call with [UserMessage]                                                                                │
│      receive AssistantMessage: "Julia is a programming language..."                                             │
│      push AssistantMessage to current_context.messages                                                          │
│      push AssistantMessage to new_messages                                                                      │
│      emit(MessageStart/End) for assistant message                                                               │
│                                                                                                                 │
│   5. check stop_reason → continue (no tools, no error)                                                          │
│                                                                                                                 │
│   6. emit(TurnEndEvent)                                                                                         │
│                                                                                                                 │
│   7. prepare_next_turn() → nothing (default)                                                                    │
│                                                                                                                 │
│   8. should_stop_after_turn() → false (default)                                                                 │
│                                                                                                                 │
│   9. pending_messages = get_steering() = []  ← No steering messages                                             │
│                                                                                                                 │
│   10. follow_up_messages = get_follow_up() = []                                                                 │
│                                                                                                                 │
│   11. break  ← Exit main loop                                                                                   │
│                                                                                                                 │
│   12. emit(AgentEndEvent)                                                                                       │
│                                                                                                                 │
│   ┌───────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│   │ Current context.messages:                                                                                 │ │
│   │   [UserMessage("What is Julia?"), AssistantMessage("Julia is...")]                                        │ │
│   │                                                                                                           │ │
│   │ steering_queue: []                                                                                        │ │
│   │ follow_up_queue: []                                                                                       │ │
│   └───────────────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│                                                                                                                 │
│   LLM SEES (convert_to_llm() filters):                                                                        │ │
│   ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐         │
│   │ Messages passed to LLM API:                                                                     │         │
│   │   [UserMessage("What is Julia?"), AssistantMessage("Julia is...")]                             │         │
│   └─────────────────────────────────────────────────────────────────────────────────────────────────┘         │
│                                                                                                                 │
│   TURN #2: User asks "How does it work?"                                                                        │
│   ─────────────────────────────────────────                                                                     │
│   1. Agent.prompt("How does it work?")                                                                          │
│      normalizePrompt() → [UserMessage("How does it work?")]                                                     │
│      runPromptMessages()                                                                                        │
│                                                                                                                 │
│   2. runAgentLoop()                                                                                             │
│      new_messages = [UserMessage("How does it work?")]                                                          │
│      current_context.messages = vcat([...previous..., UserMessage("How does it work?")])                        │
│         │                                                                                                       │
│         └─→ User message added (context preserved from Turn #1)                                                │
│      emit(AgentStartEvent), emit(TurnStartEvent)                                                                │
│      emit(MessageStart/End) for user message                                                                    │
│                                                                                                                 │
│   3. runLoop()                                                                                                  │
│      pending_messages = get_steering() = []                                                                     │
│                                                                                                                 │
│   4. streamAssistantResponse()                                                                                  │
│      convert_to_llm([UserMsg1, AssistantMsg1, UserMsg2]) → Message[]                                            │
│      LLM call with FULL conversation history (context preserved!)                                               │
│      receive AssistantMessage: "It works by..."                                                                 │
│      push AssistantMessage to current_context.messages                                                          │
│      push AssistantMessage to new_messages                                                                      │
│                                                                                                                 │
│   5. emit(TurnEndEvent), emit(AgentEndEvent)                                                                    │
│                                                                                                                 │
│   ┌───────────────────────────────────────────────────────────────────────────────────────────────────────────┐ │
│   │ Current context.messages:                                                                                 │ │
│   │   [UserMsg1, AssistantMsg1, UserMsg2, AssistantMsg2]                                                      │ │
│   └───────────────────────────────────────────────────────────────────────────────────────────────────────────┘ │
│                                                                                                                 │
│   LLM SEES:                                                                                                     │
│   ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐         │
│   │ Messages passed to LLM API:                                                                     │         │
│   │   [UserMessage("What is Julia?"),                                                              │         │
│   │    AssistantMessage("Julia is..."),                                                            │         │
│   │    UserMessage("How does it work?"),                                                           │         │
│   │    AssistantMessage("It works by...")]                                                         │         │
│   └─────────────────────────────────────────────────────────────────────────────────────────────────┘         │
│                                                                                                                 │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ STEERING MESSAGES                                                                                               │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│                                                                                                                 │
│   What is a steering message?                                                                                   │
│   • A message (any AgentMessage type) injected via: `agent.steer(message)`                                      │
│   • Goes into the steering queue, not immediately to context.messages                                          │
│                                                                                                                 │
│   How is it created?                                                                                            │
│   • User code calls: agent.steer(UserMessage("..."))                                                           │
│   • Or: agent.steer(AssistantMessage("..."))                                                                   │
│   • Or any other AgentMessage subtype                                                                          │
│                                                                                                                 │
│   When is it processed?                                                                                         │
│   • At the START of the next loop iteration (line 194-202 in agent_loop.jl)                                     │
│   • AFTER the previous assistant turn completes                                                                 │
│   • BEFORE the next assistant response is streamed                                                             │
│                                                                                                                 │
│   Why use steering?                                                                                             │
│   Use case 1: Tool execution result injection                                                                  │
│     - Agent calls a tool (e.g., read_file, bash)                                                               │
│     - Tool returns result                                                                                      │
│     - You want to inject a follow-up question based on the result                                              │
│     - agent.steer(UserMessage("Based on the file, what should we do next?"))                                   │
│                                                                                                                 │
│   Use case 2: Multi-turn conversation without user input                                                       │
│     - Agent responds to user                                                                                   │
│     - Before user types again, you want to inject a system message                                             │
│     - agent.steer(BashExecutionMessage(...)) or custom message                                                │
│     - This continues the conversation automatically                                                             │
│                                                                                                                 │
│   Use case 3: Branch navigation recovery                                                                       │
│     - User navigates between conversation branches                                                             │
│     - After switching branches, you want to inject a context message                                           │
│     - agent.steer(BranchSummaryMessage(...))                                                                   │
│     - The agent can then continue from the new branch context                                                  │
│                                                                                                                 │
│   Use case 4: Compaction summary injection                                                                     │
│     - Conversation history is compacted                                                                        │
│     - After compaction, inject summary message                                                                 │
│     - agent.steer(CompactionSummaryMessage(...))                                                               │
│     - Agent knows old history was summarized                                                                    │
│                                                                                                                 │
│   Example:                                                                                                      │
│     agent.steer(UserMessage("Follow-up question here"))                                                        │
│     # This will be processed in the next loop iteration,                                                       │
│     # appearing in context.messages before the next LLM call                                                  │
│                                                                                                                 │
│   The LLM sees:                                                                                                 │
│   ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐         │
│   │ All messages become Message[] via convert_to_llm():                                           │         │
│   │   [UserMessage(...), AssistantMessage(...), UserMessage(from_steer), ...]                    │         │
│   │                                                                                                 │         │
│   │ The LLM cannot tell which came from Agent.prompt() vs agent.steer()                           │         │
│   └─────────────────────────────────────────────────────────────────────────────────────────────────┘         │
│                                                                                                                 │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ LLM PROCESSING: How LLM sees messages                                                                          │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│                                                                                                                 │
│   The LLM NEVER sees "user message" vs "steering message" - it only sees Message types:                        │
│                                                                                                                 │
│   ┌─────────────────────────────────────────────────────────────────────────────────────────────────┐         │
│   │ convert_to_llm() transforms ALL AgentMessages to Message[]:                                   │         │
│   │                                                                                                 │         │
│   │   UserMessage("user")             → UserMessage (for LLM)                                     │         │
│   │   Steering UserMessage("user")    → UserMessage (for LLM) ← Same!                            │         │
│   │   AssistantMessage("assistant")   → AssistantMessage (for LLM)                                │         │
│   │   ToolResultMessage("toolResult") → ToolResultMessage (for LLM)                              │         │
│   │                                                                                                 │         │
│   │   BranchSummaryMessage            → UserMessage (wrapped in summary tags)                    │         │
│   │   CompactionSummaryMessage        → UserMessage (wrapped in summary tags)                    │         │
│   │   BashExecutionMessage            → UserMessage (if not excluded)                            │         │
│   │   CustomMessage                   → UserMessage                                               │         │
│   └─────────────────────────────────────────────────────────────────────────────────────────────────┘         │
│                                                                                                                 │
│   The difference is ONLY in HOW messages enter the system:                                                     │
│   • User messages:  Agent.prompt() → vcat() → context.messages (direct)                                        │
│   • Steering:       agent.steer() → queue → loop → context.messages (indirect)                                 │
│                                                                                                                 │
│   At LLM level: BOTH become UserMessage in the conversation!                                                   │
│                                                                                                                 │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘

┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐
│ KEY INSIGHTS                                                                                                    │
├─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┤
│                                                                                                                 │
│   1. User prompts go DIRECTLY to context.messages via vcat() in runAgentLoop()                                  │
│                                                                                                                 │
│   2. Steering queue is for messages injected via agent.steer() AFTER a turn finishes                             │
│      This allows continuing conversation without calling Agent.prompt() again                                   │
│                                                                                                                 │
│   3. Context is preserved across turns - context.messages grows with each turn                                  │
│      LLM sees the full conversation history                                                                     │
│                                                                                                                 │
│   4. At LLM level, ALL messages become Message types (UserMessage/AssistantMessage/ToolResultMessage)          │
│      The "steering" vs "user" distinction is just a control mechanism, not a message type                       │
│                                                                                                                 │
│   5. New turn is triggered by:                                                                                  │
│      - New Agent.prompt() call (adds user messages)                                                             │
│      - Steering messages (adds steering messages)                                                               │
│      - Follow-up messages (adds follow-up messages)                                                             │
│                                                                                                                 │
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘