9.3 KiB
LLMMCTS
LLMMCTS implements Monte Carlo Tree Search (MCTS) for Large Language Model (LLM) planning tasks.
Why LLM + MCTS?
Integrating LLMs with MCTS creates a powerful planning system that combines:
- LLM reasoning: Understand complex tasks, generate creative solutions, and reason about states
- MCTS optimization: Systematically explore solution spaces and find optimal trajectories
The Sparse Reward Problem in Traditional MCTS
In traditional reinforcement learning and MCTS, sparse rewards are a fundamental challenge that severely limits performance:
| Challenge | Traditional MCTS | LLM-MCTS Integration |
|---|---|---|
| Reward signal | Only at terminal states (e.g., win/loss) | Pseudo-rewards at every state via LLM |
| Value estimation | Relies on Monte Carlo sampling (high variance, slow convergence) | LLM provides progress value as heuristic |
| Sample efficiency | Low—requires many samples to discover reward | High—LLM guides search toward promising regions |
| Exploration strategy | Blind exploration until reward discovered | LLM suggests promising actions to try |
Why Sparse Rewards Are Problematic
-
High variance in value estimates — With sparse rewards, Monte Carlo estimates have high variance because few samples contribute to each node's value
-
Slow learning — Rewards must propagate backward through many layers before affecting early decisions, requiring many iterations
-
Poor exploration — Without intermediate signals, MCTS explores randomly until it偶然 discovers a reward, wasting computation
-
Local optima — Without guidance, MCTS may get stuck in suboptimal regions of the search space
How LLM-MCTS Solves This
Progress Value vs. Reward:
progressvalue— LLM's estimate of how close we are to solving the task (pseudo-reward)statevalue— Actual cumulative reward from Monte Carlo simulationsreward— Immediate reward from environment (may be sparse, only at terminal states)
The three-tier value system:
# LLM provides progress value at every node (dense, fast guidance)
progressvalue = llm_estimate(state) # Heuristic, available immediately
# MCTS computes statevalue via simulation (sparse but accurate)
statevalue = monte_carlo_average(simulations) # Accurate but expensive
# Environment provides immediate reward (may be sparse)
reward = environment_reward(state) # Only at terminal states
How MCTS uses these values together:
| Phase | Which value used | Why |
|---|---|---|
| Node selection (UCT) | progressvalue + statevalue |
Dense guidance for fast exploration |
| Simulation | statevalue + reward |
Accurate long-term estimates |
| Backpropagation | reward |
Ground truth updates |
Why this matters: LLMs provide dense progressvalue guidance at every node, allowing MCTS to focus computation on promising trajectories. The simulation phase confirms these estimates with accurate statevalue computed via Monte Carlo. This combination solves the sparse reward problem while maintaining accuracy.
Benefits of LLM-MCTS Integration
| Benefit | Description |
|---|---|
| Overcomes sparse rewards | LLM provides progressvalue at every node, enabling fast learning without waiting for terminal rewards |
| Faster convergence | Dense guidance from LLM reduces sample complexity by 5-10x compared to pure Monte Carlo |
| Better than pure LLM | MCTS systematically compares multiple LLM-generated trajectories, avoiding local optima |
| Better than pure planning | LLM handles complex reasoning and novel state generation that pure planners cannot |
| Uncertainty quantification | Visit counts in MCTS nodes reflect confidence in LLM's progress estimates |
| Configurable depth | MCTS depth controls planning horizon; LLM handles long-term reasoning at each step |
| Parallel exploration | MCTS naturally supports parallel simulation; LLM generates diverse candidate actions |
Overview
Key Features
- UCT-based node selection: Uses Upper Confidence Bound for Trees to balance exploration/exploitation
- Configurable expansion: Parallel or sequential child node generation
- Simulation with depth control: Rollouts with configurable maximum depth
- Reward discounting: Backpropagation with configurable future reward decay
- Multithreading support: Parallel simulation phase for improved performance
Integration with LLMs
This package is designed to work with LLMs as the state transition engine:
# LLM-based transition function
function llm_transition(state::Dict, args::NamedTuple)
# LLM generates next thought/action based on current state
response = llm_call(state[:thoughtHistory], args.prompt)
# Parse LLM output into new state
return Dict(
:newNodeKey => generate_key(),
:newstate => update_state(state, response),
:progressvalue => estimate_value(response)
)
end
result = runMCTS(initial_state, llm_transition, args)
Installation
using Pkg
Pkg.add("LLMMCTS")
Usage
Basic Example
using LLMMCTS
# Define transition function
function transition(state::Dict, args::NamedTuple)
# Your transition logic here
return Dict(:newNodeKey => "child_1", :newstate => new_state, :progressvalue => 5)
end
# Define transition arguments
transition_args = (param1 = "value1", param2 = "value2")
# Run MCTS
result = runMCTS(
initialstate,
transition,
transition_args;
maxiterations = 10,
explorationweight = 1.0,
maxSimulationDepth = 3
)
# Access results
root = result.root
best_next_state = result.bestNextState
best_terminal_state = result.bestTerminalState
high_value_states = result.highValueStateList
Advanced Usage
# With custom parameters
result = runMCTS(
initialState,
transition_func,
transition_args;
horizontalSampleExpansionPhase = 5, # More children during expansion
horizontalSampleSimulationPhase = 3, # Sample 3 children during simulation
maxSimulationDepth = 5, # Deeper search
maxiterations = 50, # More iterations
explorationweight = 2.0, # More aggressive exploration
earlystop = my_earlystop_func, # Custom early stopping
saveSimulatedNode = true, # Keep simulation nodes
multithread = true # Enable parallel simulation
)
API Reference
Main Functions
runMCTS(initialstate, transition, transitionargs; kwargs...)
Search for the best action to take for a given state and task.
Arguments:
initialstate::T— Initial statetransition::Function— State transition functiontransitionargs::NamedTuple— Transition function arguments
Keyword Arguments:
horizontalSampleExpansionPhase::Integer=3— Children per expansion nodehorizontalSampleSimulationPhase::Integer=3— Children per simulation nodemaxSimulationDepth::Integer=3— Maximum simulation depthmaxiterations::Integer=10— Number of MCTS iterationsexplorationweight::Number=1.0— Exploration weight (1.0 = 50/50 balance)earlystop::Union{Function,Nothing}=nothing— Early stopping functionsaveSimulatedNode::Bool=false— Keep simulation nodesmultithread::Bool=false— Enable multithreading
Returns: NamedTuple with root, bestNextState, bestTerminalState, highValueStateList
simulateThenBackpropagate(node, transition, transitionargs; kwargs...)
Run simulation from a node and backpropagate the reward.
Utility Functions
UCTselect(node, w)— Select node using UCT scoredictify(x; keytype=Any, stringkey=false)— Convert JSON.Object/OrderedDict to plain Dict
MCTS Node Structure
MCTSNode(
nodekey::String,
state::Dict,
visits::Integer,
progressvalue::Number,
statevalue::Number,
reward::Number,
isterminal::Bool,
parent::Union{MCTSNode, Nothing},
children::Dict{String, MCTSNode},
etc::Dict{String, Any}
)
Understanding progressvalue, statevalue, and reward
| Field | Source | Purpose |
|---|---|---|
progressvalue |
LLM heuristic | Estimate of how close we are to solving; used for fast node selection |
statevalue |
Monte Carlo simulation | Actual cumulative reward from simulations; accurate but expensive to compute |
reward |
Environment | Immediate reward (may be sparse, only at terminal states) |
Why this matters: In traditional MCTS, sparse rewards force extensive exploration. Here, LLM provides dense progressvalue guidance at every node, while statevalue (computed via simulation) provides accurate long-term estimates. MCTS balances both via UCT:
- High
progressvalue→ explored early (fast guidance) - High
statevalue→ exploited once confirmed (accurate value)
Contributing
Contributions are welcome! Please open issues for bugs or feature requests, and submit PRs for improvements.
License
MIT License — see LICENSE for details.
Author
narawat lamaiin narawat@outlook.com