Adding memory to search instead of sampling in reward maximization tasks [R]
FLEET is an algorithm designed to improve Best-of-N generation by integrating external rewards with specific tokens and utilizing Monte Carlo Tree Search (MCTS) to refine logits in subsequent runs. The algorithm's authors have released a preprint (2609.27657), a Huggingface page, and a GitHub repository containing experiments, examples, and a Python package. Further details on MCTS modifications and parameter tuning for various models and tasks are available.
This paper introduces FLEET, an algorithm that enhances Best-of-N generation by using MCTS to adjust logits, unlike previous methods that relied on sampling.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
IngestedOffset at this time: UTC+0Oct 2, 2026, 22:00 UTC
- Ingested
- Oct 2, 2026, 22:00
- Source type
- Dev community
Full text isn't available here.
Read at source →