About this project
LightZero is an open-source algorithm toolkit from OpenDILab that combines Monte Carlo Tree Search (MCTS) with deep reinforcement learning in a single PyTorch-based framework. It was presented as a Spotlight paper at the NeurIPS 2023 Datasets and Benchmarks Track and is positioned as a unified benchmark for MCTS in general sequential decision scenarios.
The project emphasizes three properties: being lightweight, efficient, and easy to understand. It integrates multiple MCTS algorithm families so that decision-making problems with different attributes can be handled within one framework. For computational efficiency, it uses mixed heterogeneous computing for the most time-consuming parts of MCTS, with tree implementations in both Python (ptree) and C++ (ctree). Documentation includes algorithm framework diagrams, function call graphs, and network structure diagrams to help users locate critical code and compare algorithms under a shared paradigm.
Implemented algorithms include AlphaZero, MuZero, Sampled MuZero, Stochastic MuZero, EfficientZero, Sampled EfficientZero, Gumbel MuZero, ReZero, and UniZero. Supported environments span board games (TicTacToe, Gomoku, Connect4), 2048, classic control tasks (CartPole, Pendulum, LunarLander, BipedalWalker), Atari, DeepMind Control, MuJoCo, MiniGrid, Bsuite, Memory, SumToThree, and MetaDrive. A compatibility table in the README marks which algorithm-environment combinations are finished and tested, which are work in progress, and which are unsupported.
The framework is organized around three core modules: Model (network structure and forward propagation), Policy (learning, collecting, and evaluation processes), and MCTS (search tree structure and its interaction with the policy). Installation is via cloning the repository and running pip install -e ., with support currently limited to Linux and macOS; a Dockerfile based on Ubuntu 20.04 with Python 3.8 is also provided. Quick-start examples train MuZero on CartPole, Pong, and TicTacToe, and UniZero on Pong. An experimental Ray-based async segment training pipeline for Atari is documented, with the README reporting a same-window comparison against a synchronous baseline.
Benchmarks are published for AlphaZero and MuZero on board games, for MuZero variants on Atari titles such as Pong, Qbert, and MsPacman, for Sampled EfficientZero with factored and Gaussian policy representations on continuous-control tasks, for Gumbel MuZero under different simulation budgets, and for Stochastic MuZero on 2048 with varying chance levels. The repository also maintains an Awesome-MCTS section collecting foundational and recent papers on AlphaGo, MuZero, MCTS analysis, and applications, plus tutorials on customizing environments and algorithms, configuration files, logging, and loss landscape visualization.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.