SEARCH · Engineering Papers
Results for “multi-objective hyperparameter optimization”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Multi-Objective Hyperparameter Optimization for Spiking Neural Network Neuroevolution
Neuroevolution has had significant success over recent years, but there has been relatively little work applying neuroevolution approaches to spiking neural networks (SNNs). SNNs are a type of neural network that includes temporal processing component, are not easily trained using other methods, and can be deployed into energy-efficient neuromorphic hardware. In this work, we investigate two evolutionary approaches for training SNNs. We explore the impact of the hyperparameters of the evolutionary approaches, including tournament size, population size, and representation type, on the performance of the algorithms. We present a multi-objective Bayesian-based hyperparameter optimization approach to tune the hyperparameters to produce the most accurate and smallest SNNs. We show that the hyperparameters can significantly affect the performance of these algorithms. We also perform sensitivity analysis and demonstrate that every hyperparameter value has the potential to perform well, assuming other hyperparameter values are set correctly.
HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models
As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.
A general Bayesian algorithm for the autonomous alignment of beamlines
Autonomous methods to align beamlines can decrease the amount of time spent on diagnostics, and also uncover better global optima leading to better beam quality. The alignment of these beamlines is a high-dimensional expensive-to-sample optimization problem involving the simultaneous treatment of many optical elements with correlated and nonlinear dynamics. Bayesian optimization is a strategy of efficient global optimization that has proved successful in similar regimes in a wide variety of beamline alignment applications, though it has typically been implemented for particular beamlines and optimization tasks. In this paper, we present a basic formulation of Bayesian inference and Gaussian process models as they relate to multi-objective Bayesian optimization, as well as the practical challenges presented by beamline alignment. We show that the same general implementation of Bayesian optimization with special consideration for beamline alignment can quickly learn the dynamics of particular beamlines in an online fashion through hyperparameter fitting with no prior information. We present the implementation of a concise software framework for beamline alignment and test it on four different optimization problems for experiments on X-ray beamlines at the National Synchrotron Light Source II and the Advanced Light Source, and an electron beam at the Accelerator Test Facility, along with benchmarking on a simulated digital twin. We discuss new applications of the framework, and the potential for a unified approach to beamline alignment at synchrotron facilities.
HPS-RL: Hyperparameter tuning for deep RL applications (HPS-RL) v1
Genetic Algorithms meets Deep RL for Hyperparameters Hyperparameter optimization and architecture search can easily become cumbersome and finding the right hyperparameters can seriously impact the robustness of the deep RL application being developed. We use genetic algorithms to evolve optimum deep RL architectures in a scalable manner. HPS-RL is designed to work with multiple gym enviornments, allow users to test their own optimization functions and tune multi-objective parameters in multiple deep RL algorithms. HPS-RL uses multi-threading and is being extended with mpipy for distributed processing on HPC. https://arxiv.org/abs/2201.11182
Numerical modeling based machine learning approach for the optimization of falling - film evaporator in thermal desalination application
Scale formation that drastically increases thermal resistance and reduces freshwater production remains a critical challenge in thermal desalination. Novel designs of falling film evaporator and optimal operating condition hold great promise to mitigate scale formation, and increase heat transfer performance and fresh water production. In this work, CFD simulation based machine learning and multi-objective optimization are performed to identify optimal conditions and tube arrangement for evaporator. Non-dominated sorting genetic algorithm is adopted to determine and analyze the optimal pareto front for multiple objectives in desalination criteria. The errors of training, validation, and testing set are computed to identify an optimal hyperparameter set. For performance ratio, fouling resistance, and water production rate, the average relative error is 2.26%, 3.67%, and 3.24%. At pareto front, both performance ratio and water production rate increase at high temperature with fouling resistance (thermal resistance of the fouling layer) increasing as well. Tradeoffs between mitigating scale formation and enhancing desalination performance are evaluated in optimizations for different objectives. Finally, potential optima are identified and can be applied as guidelines to determine evaporator design and system operating conditions.