DOE OSTI · code-153657
Trajectory Balance with Asynchrony
Abstract
Finetune large language models (LLMs) quicker via parallelization. Specifically, this implements asynchronous reinforcement learning with a trajectory balance objective function
Keep this discovery
Explore connections, maps & timelines
Bartoldson, BrianR [Lawrence Livermore National Laboratory (LLNL), Livermore, CA (United States)]. 2025-03-21. Trajectory Balance with Asynchrony. https://doi.org/10.11578/dc.20250410.2
Cite the original work for its findings. Save a collection to share your selection of sources.