Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Parallel Variable Population Multi-Objective Optimizer (pvpmoo) v1.0

This is a parallel variable population multi-objective optimizer with an adaptive unified differential evolution algorithm or a genetic algorithm. It can also be used for single objective optimization. Some features of this code include: 1) The population size varies from generation to generation to save the total # of objective function evaluations. 2) The population is uniformly distributed to a number of parallel processors for simultaneous objective function evaluation. 3) The objective function evaluation can be attained from an external simulation program with control variables in its input file and objectives calculated from its output files. 4) The optimizer includes an adaptive unified differential evolution algorithm and a real value genetic algorithm. The parameters in the unified differential evolution algorithm can be chosen to attain any mutation schemes in the published literature.

Qiang, Ji↗

Parallel derivative-free optimization for simulation-based design of behind-the-meter energy systems

In this work, the integrated design and dispatch of behind-the-meter or distributed resources (e.g. stationary battery storage and solar PV generation) is considered. A simulation-based framework is employed, generating high-fidelity results with closed-loop predictive control at a fine resolution, at the expense of high computational cost (several minutes to a few hours per design point). To address this challenge, parallel derivative-free design methods are considered. Four methods are compared, including state-of-the-art surrogate-based methods (Radial-Basis Functions and Gaussian processes) and sampling strategies, an evolutionary-based method, and a simple sequential grid refinement method. As a case study, two types of design problem with increasing complexity are considered, namely, the design of behind-the-meter resources (three design variables) and the inclusion of grid capacity (four design variables). The second yields a constrained design problem for which violations can only be determined after solving the computationally expensive simulation. For the three-dimensional case, all methods present a good performance, achieving a solution within 1% of the optimum after the first iteration, with the sequential grid refinement exhibiting the fastest convergence and achieving the best final objective value. This indicates that the parallel evaluation of multiple sampling points may be more important than the choice of method for small decision spaces. For the four-dimensional constrained case, the Genetic Algorithm presents the best tradeoff between performance and computational effort, while the rough objective function terrain generated by constraint violation penalties reduces the performance of surrogate-based methods. Contour plots with flat regions indicate flexibility in the optimal design and highlight the importance of characterizing the solution space.

24 POWER TRANSMISSION AND DISTRIBUTION↗

HBMax: Optimizing Memory Efficiency for Parallel Influence Maximization on Multicore Architectures

The goal of influence maximization is to select k most-influential vertices or seeds in a network, where influence is defined by a given diffusion process. The problem has a number of important applications such as viral marketing, information spread, and epidemic control. Although computing optimal seed set is NP-Hard, due to the submodular nature of the problem efficient approximation algorithms exist. However, even state-of-the-art parallel implementations are limited by a sampling step that incurs large memory footprints. This in turn limits the problem size reach and approximation quality. In this work, we study the memory footprint of the sampling process collecting reverse reachability information in the IMM algorithm over large real-world social networks. We present an adaptive and memory-efficient optimization approach for a state-of-the-art multi-threaded parallel influence maximization algorithm. Our approach,HuffMax, uses a portion of the reverse reachable (RR) sets collected by the algorithm to learn the characteristics of the graph. Then, it compresses the intermediate reverse reachability information with Huffman coding, and queries directly on the compressed data to preserve the memory savings obtained through compression. We also propose an efficient sampling strategy based on the distribution of RR sets, which can further reduce the computation time for typical social networks with long-tail distributions. Considering a NUMA architecture, we scale up our solution on 128-core CPUs and reduce the memory footprint by up to 45.7% with negligible time overhead (or even faster) and without perceivable loss of accuracy.

Chen, Xinyu↗

HYPPO: A Surrogate-Based Multi-Level Parallelism Tool for Hyperparameter Optimization

We present a new software, HYPPO, that enables the automatic tuning of hyperparameters of various deep learning (DL) models. Unlike other hyperparameter optimization (HPO) methods, HYPPO uses adaptive surrogate models and directly accounts for uncertainty in model predictions to find accurate and reliable models that make robust predictions. Using asynchronous nested parallelism, we are able to significantly alleviate the computational burden of training complex architectures and quantifying the uncertainty. HYPPO is implemented in Python and can be used with both TensorFlow and PyTorch libraries. We demonstrate various software features on time-series prediction and image classification problems as well as a scientific application in computed tomography image reconstruction. Finally, we show that (1) we can reduce by an order of magnitude the number of evaluations necessary to find the most optimal region in the hyperparameter space and (2) we can reduce by two orders of magnitude the throughput for such HPO process to complete.

adaptation models↗

OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model

With the sharp increasing volume of user data, Deep Learning Recommendation Model (DLRM) becomes an indispensable infrastructure in large technology companies. However, large-scale DLRM on the multi-GPU platform is still inefficient due to unbalanced workload partitioning and intensive inter-GPU communication. To this end, we propose OPER, an OPtimality guided Embedding table placement for large-scale Recommendation model training and inference. OPER explores the potential of mitigating remote memory access latency in DLRM through fine-grained embedding table placement. Specifically, OPER proposes a theoretical modeling that builds up the relationship between EMT placement and the embedding communication latency in both training and inference. OPER proves the NP hardness of finding the optimal embedding table placement and proposes a heuristic algorithm that yields near optimal placement. OPER implements a SHMEM-based embedding table training system and a unified embedding index mapping to support fine-grained embedding table sharding and placement. Comprehensive experiments reveal that OPER achieves on average 3.4× and 5.1× speedup on training and inference respectively over state-of-the-art DLRM frameworks.

Wang, Zheng↗

Process–Property–Performance Mapping of Additively Manufactured 316H Stainless Steel Components

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development of advanced materials and components fabricated via additive manufacturing, and is using laser powder bed fusion (LPBF) of 316H stainless steel as an initial case study. In the previous fiscal year, miniature high-throughput specimens were printed on multiple LPBF systems to provide initial processing windows to minimize porosity and limit epitaxial grain growth during prints. This fiscal year, scaled builds were completed on three different LPBF systems at ORNL: a GE Concept Laser M2, a Renishaw AM400, and an EOS M290. Builds on the Concept Laser were conducted on multiple powder lots and processing parameter ranges to provide microstructure effects on time-independent and time-dependent mechanical properties. Builds on the Renishaw were produced using Oak Ridge National Laboratory (ORNL)-optimized printing parameters and Argonne National Laboratory (ANL)-optimized printing parameters to compare outcomes of parallel process optimization efforts at different national laboratories on the same LPBF system. Similarly, the build completed on the EOS M290 replicated the processing parameters of builds completed at Los Alamos National Laboratory (LANL). Optical microscopy and electron backscatter diffraction characterization was completed on all builds. In addition to the general round robin characterization, this work-package generated time-independent data, including tensile and fracture toughness test data on scaled Concept Laser builds as a function of processing parameters and post-build heat treatment. This analysis is complimentary to work in parallel work packages aiming to establish heat treatment and processing effects on time-dependent properties. It was found that although the stress-relief heat treatment provides the highest strength at lower-temperatures, tensile strength begins to converge at higher temperatures regardless of heat treatment condition. In addition, the more rigorous solution annealing and hot-isostatic pressing post-build heat treatments result in higher fracture toughness than the stress-relieved condition. The root-causes of the lower fracture toughness of the stress-relieved LPBF 316H material was informed via a stress-relief optimization study on a scaled concept laser print, where it was found that although dislocation recovery was largely complete after only a couple hours at 650°C, the extended hold of the current 24h heat treatment employed on scaled builds likely caused increased carbide volume fractions along the LPBF 316H grain boundaries, thereby deteriorating crack propagation resistance. This trend was seen to become more deleterious with additional increases of stress-relief temperature to 750°C or 850°C. These results have helped inform a new optimal stress-relief annealing condition for LPBF 316H for future campaign testing (650°C for 2h).

36 MATERIALS SCIENCE↗

Process–Property–Performance Mapping of Additively Manufactured 316H Stainless Steel Components

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development of advanced materials and components fabricated via additive manufacturing, and is using laser powder bed fusion (LPBF) of 316H stainless steel as an initial case study. In the previous fiscal year, miniature high-throughput specimens were printed on multiple LPBF systems to provide initial processing windows to minimize porosity and limit epitaxial grain growth during prints. This fiscal year, scaled builds were completed on three different LPBF systems at ORNL: a GE Concept Laser M2, a Renishaw AM400, and an EOS M290. Builds on the Concept Laser were conducted on multiple powder lots and processing parameter ranges to provide microstructure effects on time-independent and time-dependent mechanical properties. Builds on the Renishaw were produced using Oak Ridge National Laboratory (ORNL)-optimized printing parameters and Argonne National Laboratory (ANL)-optimized printing parameters to compare outcomes of parallel process optimization efforts at different national laboratories on the same LPBF system. Similarly, the build completed on the EOS M290 replicated the processing parameters of builds completed at Los Alamos National Laboratory (LANL). Optical microscopy and electron backscatter diffraction characterization was completed on all builds. In addition to the general round robin characterization, this work-package generated time-independent data, including tensile and fracture toughness test data on scaled Concept Laser builds as a function of processing parameters and post-build heat treatment. This analysis is complimentary to work in parallel work packages aiming to establish heat treatment and processing effects on time-dependent properties. It was found that although the stress-relief heat treatment provides the highest strength at lower-temperatures, tensile strength begins to converge at higher temperatures regardless of heat treatment condition. In addition, the more rigorous solution annealing and hot-isostatic pressing post-build heat treatments result in higher fracture toughness than the stress-relieved condition. The root-causes of the lower fracture toughness of the stress-relieved LPBF 316H material was informed via a stress-relief optimization study on a scaled concept laser print, where it was found that although dislocation recovery was largely complete after only a couple hours at 650°C, the extended hold of the current 24h heat treatment employed on scaled builds likely caused increased carbide volume fractions along the LPBF 316H grain boundaries, thereby deteriorating crack propagation resistance. This trend was seen to become more deleterious with additional increases of stress-relief temperature to 750°C or 850°C. These results have helped inform a new optimal stress-relief annealing condition for LPBF 316H for future campaign testing (650°C for 2h).

36 MATERIALS SCIENCE↗

VAN-DAMME: GPU-accelerated and symmetry-assisted quantum optimal control of multi-qubit systems

We present an open-source software package, VAN-DAMME (Versatile Approaches to Numerically Design, Accelerate, and Manipulate Magnetic Excitations), for massively-parallelized quantum optimal control (QOC) calculations of multi-qubit systems. To enable large QOC calculations, the VAN-DAMME software package utilizes symmetry-based techniques with custom GPU-enhanced algorithms. This combined approach allows for the simultaneous computation of hundreds of matrix exponential propagators that efficiently leverage the intra-GPU parallelism found in high-performance GPUs. In addition, to maximize the computational efficiency of the VAN-DAMME code, we carried out several extensive tests on data layout, computational complexity, memory requirements, and performance. These extensive analyses allowed us to develop computationally efficient approaches for evaluating complex-valued matrix exponential propagators based on Padé approximants. To assess the computational performance of our GPU-accelerated VAN-DAMME code, we carried out QOC calculations of systems containing 10 - 15 qubits, which showed that our GPU implementation is 18.4× faster than the corresponding CPU implementation. Our GPU-accelerated enhancements allow efficient calculations of multi-qubit systems, which can be used for the efficient implementation of QOC applications across multiple domains.

97 MATHEMATICS AND COMPUTING↗

PARMOO

ParMOO is a Python library for solving multiobjective simulation optimization problems, while exploiting problem structure. ParMOO stands for "parallel multiobjective optimization".

WILD, STEFAN↗

OpenGraphGym: A Parallel Reinforcement Learning Framework for Graph Optimization Problems

This paper presents an open-source, parallel AI environment (named OpenGraphGym) to facilitate the application of reinforcement learning (RL) algorithms to address combinatorial graph optimization problems. This environment incorporates a basic deep reinforcement learning method, and several graph embeddings to capture graph features, it also allows users to rapidly plug in and test new RL algorithms and graph embeddings for graph optimization problems. This new open-source RL framework is targeted at achieving both high performance and high quality of the computed graph solutions. This RL framework forms the foundation of several ongoing research directions, including 1) benchmark works on different RL algorithms and embedding methods for classic graph problems; 2) advanced parallel strategies for extreme-scale graph computations, as well as 3) performance evaluation on real-world graph solutions.

Zheng, Weijian↗

Parallel Time Integration for Constrained Optimization

The number of transistors in an average processor continues to increase, but individual clock speeds have plateaued. Those transistors are instead going into additional cores, increasing the number of different things that a processor can do at once and placing an emphasis on parallel computation. Many problems in scientific computing follow a time-evolution model, and it can be difficult to solve such problems in parallel across the temporal domain. The Multi-Grid Reduction In Time (MGRIT) algorithm, developed at Lawrence Livermore National Laboratory (LLNL), solves differential equations with a method designed specifically to take advantage of extreme numbers of processors by parallelizing across time. The Tri-diagonal MGRIT (TriMGRIT) algorithm, also developed at LLNL, is a generalization of MGRIT which enables parallel-in-time solving of a greater number of problems. Constrained optimization problems, in particular, may be solved in parallel using TriMGRIT. These consist of choosing a control function such that an objective functional is minimized, constrained by a differential-equation. We consider two such problems: applying torque to a pendulum to bring it to a gentle stop and moving a crowd of people from one distribution into another. We also perform some miscellaneous theoretical and practical research, including investigating the use of a line-search subroutine to refine intermediate TriMGRIT results and preliminary work on strategies for choosing operators for TriMGRIT to use.

97 MATHEMATICS AND COMPUTING↗

1st Computational Physics School for Fusion Research (2019 CPS-FR)

The rising number of applications of machine learning and computational statistics in fusion energy research requires flexibility in adopting a growing variety of tools. The Computational Physics School for Fusion Research (CPS-FR) aims at providing young researchers with critical skill sets to deal with modern fusion energy research challenges. The School aims at covering essentials of: Computational Statistics, Machine Learning, Deep Learning and optimization methods, Parallel Programming and HPC. As the first edition of the CPS-FR just concluded, this report highlights its main results and summarizes its contents.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

GentenMPI: Distributed Memory Sparse Tensor Decomposition

GentenMPl is a toolkit of sparse canonical polyadic (CP) tensor decomposition algorithms that is designed to run effectively on distributed-memory high-performance computers. Its use of distributed-memory parallelism enables it to efficiently decompose tensors that are too large for a single compute node's memory. GentenMPl leverages Sandia's decades-long investment in the Trilinos solver framework for much of its parallel-computation capability. Trilinos contains numerical algorithms and linear algebra classes that have been optimized for parallel simulation of complex physical phenomena. This work applies these tools to the data science problem of sparse tensor decomposition. In this report, we describe the use of Trilinos in GentenMPl, extensions needed for sparse tensor decomposition, and implementations of the CP-ALS (CP via alternating least squares) and GCP-SGD (generalized CP via stochastic gradient descent) sparse tensor decomposition algorithms. We show that GentenMPl can decompose sparse tensors of extreme size, e.g., a 12.6-terabyte tensor on 8192 computer cores. We demonstrate that the Trilinos backbone provides good strong and weak scaling of the tensor decomposition algorithms.

97 MATHEMATICS AND COMPUTING↗

Extreme-scale stochastic optimization and simulation via learning-enhanced decomposition and parallelization (Final Technical Report)

Stochastic optimization and simulation models ubiquitously arise in designing and operating complex service/engineering systems. They can be extreme in scale due to high-dimensional data and decisions, and can also involve decisions made sequentially in response to newly revealed data, both causing significant computational challenge. The objective of this research is to explore a unified framework that integrates machine learning with discrete optimization and risk-averse modeling, to improve the efficiency of decomposition paradigms for stochastic optimization and simulations at extreme scale. The models we consider represent a broad class of complex decision-making problems, where 0-1 or continuous decisions are made before and/or after knowing multiple sources of uncertainties that could be correlated. We will employ machine learning methods to dynamically decide and prioritize computational procedures, including cut generation, branching, and bounding of the optimal objective. Furthermore, the research will shed new lights on the traditional decomposition algorithms for extreme-scale computing. Deliverables of the research include new modeling and computational methods for advancing the state-of-the-art research in optimization and simulation, bringing many relevant risk-averse, data-driven optimization problems in practice within the range of tractability. Examples include distributed computing server scheduling and sensor deployment for monitoring critical infrastructures. Success in this effort will enable progress in solving multiple extreme-scale problems in the complex system design and operations arising from DoE missions in energy, environment, and national security.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Performing optimized collective operations in a irregular subcommunicator of compute nodes in a parallel computer

In a parallel computer, performing optimized collective operations in an irregular subcommunicator of compute nodes may be carried out by: identifying, within the irregular subcommunicator, regular neighborhoods of compute nodes; selecting, for each neighborhood from the compute nodes of the neighborhood, a local root node; assigning each local root node to a node of a neighborhood-wide tree topology; mapping, for each neighborhood, the compute nodes of the neighborhood to a local tree topology having, at its root, the local root node of the neighborhood; and performing a one way, rooted collective operation within the subcommunicator including: performing, in one phase, the collective operation within each neighborhood; and performing, in another phase, the collective operation amongst the local root nodes.

Davis, Kristan Suzanne D.↗

CCM vs. CRM Design Optimization of a Boost-derived Parallel Active Power Decoupler for Microinverter Applications

Single-phase inverter or rectifier systems often make use of an auxiliary active power decoupler (APD) to balance the mismatch between steady DC power and fluctuating AC power. This paper deals with efficiency and size optimization of a parallel boost-type APD circuit for PV microinverter applications. Specifically, design of an eGaNFET-based, 400 W APD circuit, employing planar inductor and operating in either continuous conduction mode (CCM) or critical conduction mode (CRM) is considered. Available design variables including inductance value, inductor core geometry, capacitor voltage, switching frequency, and modulation scheme (CCM vs. CRM) are explored to identify Pareto-optimal configurations, which can achieve low California Energy Commission (CEC) efficiency drop while also reducing the footprint area of the inductor. The theoretical study predicts that the optimal CRM design can achieve 37% reduced inductor size, while operating with similar efficiency drop, compared to the optimal CCM design. Experimental results, obtained using two separate 40 V, 400 W hardware prototypes for CCM and CRM, are presented to verify the analyses.

14 SOLAR ENERGY↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING↗