update
This commit is contained in:
@@ -13,7 +13,7 @@ MCTS is a powerful search algorithm that balances exploration and exploitation t
|
|||||||
However, in many real-world problems, **rewards are sparse**—they only come at the final state. This creates two critical problems:
|
However, in many real-world problems, **rewards are sparse**—they only come at the final state. This creates two critical problems:
|
||||||
|
|
||||||
1. **Value estimation delay** — Rewards must propagate backward through many layers before affecting early decisions
|
1. **Value estimation delay** — Rewards must propagate backward through many layers before affecting early decisions
|
||||||
2. **Exploration inefficiency** — Without intermediate signals, MCTS explores randomly until it偶然 discovers a reward
|
2. **Exploration inefficiency** — Without intermediate signals, MCTS explores randomly until it discovers a reward
|
||||||
|
|
||||||
### How LLMs Fix the Sparse Reward Problem
|
### How LLMs Fix the Sparse Reward Problem
|
||||||
|
|
||||||
@@ -138,13 +138,13 @@ function transition(state::Dict, args::NamedTuple)
|
|||||||
end
|
end
|
||||||
|
|
||||||
# Define transition arguments
|
# Define transition arguments
|
||||||
transition_args = (param1 = "value1", param2 = "value2")
|
transitionargs = (param1 = "value1", param2 = "value2")
|
||||||
|
|
||||||
# Run MCTS
|
# Run MCTS
|
||||||
result = runMCTS(
|
result = runMCTS(
|
||||||
initialstate,
|
initialstate,
|
||||||
transition,
|
transition,
|
||||||
transition_args;
|
transitionargs;
|
||||||
maxiterations = 10,
|
maxiterations = 10,
|
||||||
explorationweight = 1.0,
|
explorationweight = 1.0,
|
||||||
maxSimulationDepth = 3
|
maxSimulationDepth = 3
|
||||||
@@ -162,9 +162,9 @@ high_value_states = result.highValueStateList
|
|||||||
```julia
|
```julia
|
||||||
# With custom parameters
|
# With custom parameters
|
||||||
result = runMCTS(
|
result = runMCTS(
|
||||||
initialState,
|
initialstate,
|
||||||
transition_func,
|
transition_func,
|
||||||
transition_args;
|
transitionargs;
|
||||||
horizontalSampleExpansionPhase = 5, # More children during expansion
|
horizontalSampleExpansionPhase = 5, # More children during expansion
|
||||||
horizontalSampleSimulationPhase = 3, # Sample 3 children during simulation
|
horizontalSampleSimulationPhase = 3, # Sample 3 children during simulation
|
||||||
maxSimulationDepth = 5, # Deeper search
|
maxSimulationDepth = 5, # Deeper search
|
||||||
@@ -203,12 +203,19 @@ Search for the best action to take for a given state and task.
|
|||||||
|
|
||||||
#### `simulateThenBackpropagate(node, transition, transitionargs; kwargs...)`
|
#### `simulateThenBackpropagate(node, transition, transitionargs; kwargs...)`
|
||||||
|
|
||||||
Run simulation from a node and backpropagate the reward.
|
Run simulation from a node and backpropagate the reward. Returns `nothing`.
|
||||||
|
|
||||||
|
**Keyword Arguments:**
|
||||||
|
- `maxSimulationDepth::Integer=3` — Maximum simulation depth
|
||||||
|
- `horizontalSampleSimulationPhase::Integer=3` — Children per simulation node
|
||||||
|
- `saveSimulatedNode::Bool=false` — Keep simulation nodes
|
||||||
|
- `multithread::Bool=false` — Enable multithreading
|
||||||
|
- `highValueState` — Channel to store high-value states
|
||||||
|
|
||||||
### Utility Functions
|
### Utility Functions
|
||||||
|
|
||||||
- `UCTselect(node, w)` — Select node using UCT score
|
- `UCTselect(node, w)` — Select node using UCT score
|
||||||
- `dictify(x; keytype=Any, stringkey=false)` — Convert JSON.Object/OrderedDict to plain Dict
|
- `dictify(x; keytype=Any)` — Convert JSON.Object/OrderedDict to plain Dict
|
||||||
|
|
||||||
### MCTS Node Structure
|
### MCTS Node Structure
|
||||||
|
|
||||||
@@ -255,7 +262,7 @@ Where:
|
|||||||
| Scenario | Alternative approach |
|
| Scenario | Alternative approach |
|
||||||
|----------|---------------------|
|
|----------|---------------------|
|
||||||
| **Dense rewards available** | Use pure RL with reward shaping |
|
| **Dense rewards available** | Use pure RL with reward shaping |
|
||||||
| **Simple决策 problems** | Classical search (DFS, BFS) is sufficient |
|
| **Simple decision problems** | Classical search (DFS, BFS) is sufficient |
|
||||||
| **Real-time constraints** | LLM calls may be too slow; use pre-trained value function |
|
| **Real-time constraints** | LLM calls may be too slow; use pre-trained value function |
|
||||||
| **No LLM access** | Use pure MCTS with hand-designed heuristics |
|
| **No LLM access** | Use pure MCTS with hand-designed heuristics |
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user