Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Avoiding excess computation in asynchronous evolutionary algorithms

Abstract Asynchronous evolutionary algorithms are becoming increasingly popular as a means of making full use of many processors while solving computationally expensive search and optimization problems. These algorithms excel at keeping large clusters fully utilized, but may sometimes inefficiently sample an excess of fast‐evaluating solutions at the expense of higher‐quality, slow‐evaluating ones. We have previously introduced a steady‐state parent selection strategy, SWEET (“Selection whilE EvaluaTing”), that sometimes selects individuals that are still being evaluated and allows them to reproduce early. We perform a takeover‐time analysis that confirms that this strategy gives slow‐evaluating individuals that have higher fitnesses an increased ability to multiply in the population. We also find that SWEET appears effective at improving optimization performance on problems in which solution quality is positively correlated with evaluation time. We evaluate our approach on six simulated real‐valued optimization problems and three real‐world applications: an autonomous vehicle controller problem that involves tuning a spiking neural network and two adversarial EA problems. We further evaluate SWEET versus a basic asynchronous process in a simulated setting. We present evidence that SWEET outperforms basic asynchronous processes in a use‐case in which performance is positively correlated with evaluation time, and performs comparably (and often better) than basic asynchronous processes in several use‐cases where performance is negatively correlated with evaluation time. That said, in the cases where performance and evaluation time are negatively correlated the variance of outcomes for SWEET is notably high.

97 MATHEMATICS AND COMPUTING↗

Devastator Parallel Discrete Event Simulation Runtime (Devastator) v1.0

The Devastator runtime is a modern C++ implementation of optimistic parallel discrete event simulation methods. Devastator allows simulation application code to productively specify their component and event functionality with C++14 constructs. It utilizes GASNet-EX for distributed memory communication and includes parallel performance optimizations such as light-weight thread message queues and asynchronous GVT. Furthermore, it supports efficient event broadcasts and pause-rewind-resume functionality to support periodic load balancing and outer loop optimization algorithms.

Chan, Cy↗

SBIR Phase I Final Report, TACO: Distributed and Heterogeneous Sparse Compiler

Tensor algebra is a powerful tool for computing, but writing optimized codes that operate on sparse tensors can be very complex. This project enables a Tensor Algebra Compiler (TACO) that simplifies this task from man-years to man-days and extends TACO to support complex and large distributed systems. This report details the hypotheses, approaches used, and findings in this project.

97 MATHEMATICS AND COMPUTING↗

Cobalt-Free Cathodes for Next Generation Li-Ion Batteries

In this U.S. Department of Energy sponsored project Nexceris, in collaboration with project partners; The Ohio State University and Navitas Advanced Systems have advanced the technical maturity of a non-cobalt containing cathode for next-generation Li-ion batteries. The cathode is based on the lithium manganese nickel-titanium oxide, LiNi 0.5 Mn 1.5 TiO 4 (LNMTO) high voltage spinel. To address limitations with poor cycle and calendar life an microstucturally hierarchical LNMO/LNMTO core-shell cathode powder has been developed that enables the formation of a solid-electrolyte interface that effectively passivates the cathode surface. The microstructural enhancements of the cathode material focus on preferentially enriching the surface with titanium. In parallel, new, optimized binder and electrolyte chemistries have been incorporated to address degradation mechanisms associated with high-voltage systems. Single-layer pouch cell and large-format 2-Ah cell testing have shown that an optimized LNMO/LNMTO core-shell powder significantly improves initial cell capacity and cycle life compared to homogeneous LNMO powder. To support development and the fabrication of 2-Ah cells a novel Hybrid Alternative Wet-Chemical Synthesis (HAWCS) process have been developed. This low-cost, synthesis approach enables the excellent compositional and particle morphology control achieved with co-precipitation without the strict process controls and associated expensive process equipment.

25 ENERGY STORAGE↗

Automatic Generation of Algorithms for High-Speed Reliable Lossy Data Compression (Final Report)

Fast reliable data compression is urgently needed for many leading-edge scientific instruments and for exascale high-performance computing applications because they produce vast amounts of data at extremely high rates. The goal of this project has been to develop a framework named LC that is able to automatically generate high-speed lossless and reliable lossy compression and decompression algorithms that can be customized for different kinds of data. The resulting LC framework is freely available on GitHub. To achieve high-speed operation, LC outputs optimized and parallelized CPU and GPU implementations of the generated algorithms. To ensure the quality of lossily compressed data, LC guarantees the user-provided error bound. To be able to customize the compression algorithm to various use cases, LC can synthesize millions of different algorithms and automatically search for the one that works best for the given data. We have already employed LC to create state-of-the-art lossless and lossy compressors for scientific data as well as leading lossless compressors for images. We hope that LC and the customized, fast, reliable, and CPU/GPU-compatible compression algorithms that it can generate will greatly benefit the many scientific applications that need not only high trustworthiness but also high performance.

97 MATHEMATICS AND COMPUTING↗

Cell-free Scaled Production and Adjuvant Addition to a Recombinant Major Outer Membrane Protein from Chlamydia muridarum for Vaccine Development

Subunit vaccines offer advantages over more traditional inactivated or attenuated whole-cell-derived vaccines in safety, stability, and standard manufacturing. To achieve an effective protein-based subunit vaccine, the protein antigen often needs to adopt a native-like conformation. This is particularly important for pathogensurface antigens that are membrane-bound proteins. Cell-free methods have been successfully used to produce correctly folded functional membrane protein through the co-translation of nanolipoprotein particles (NLPs), commonly known as nanodiscs. This strategy can be used to produce subunit vaccines consisting of membrane proteins in a lipid-bound environment. However, cell-free protein production is often limited to small scale (<1 mL). The amount of protein produced in small-scale production runs is usually sufficient for biochemical and biophysical studies. However, the cell-free process needs to be scaled up, optimized, and carefully tested to obtain enough protein for vaccine studies in animal models. Other processes involved in vaccine production, such as purification, adjuvant addition, and lyophilization, need to be optimized in parallel. This paper reports the development of a scaled-up protocol to express, purify, and formulate a membrane-bound protein subunit vaccine

59 BASIC BIOLOGICAL SCIENCES↗

FLASSH 1.0: Thermal scattering law evaluation and cross section generation for reactor physics applications

The Full Law Analysis Scattering System Hub (FLASSH) is a modern, advanced code which evaluates the thermal scattering law (TSL) along with accompanying cross sections. FLASSH features generalized methods which accommodate any material structure. Historical approximations including the incoherent and cubic approximations have been removed. Instead, the latest release of FLASSH features advanced physics options including distinct corrections (1-phonon contributions) and non-cubic formulations. The non-cubic elastic and inelastic contributions are necessary to accurately evaluate 1-phonon contributions. Both non-cubic and 1-phonon calculations require high-density sampling of the various scattering directions. Optimization and parallelization of these routines were therefore necessary to produce results in a reasonable timeframe. With these notable improvements to the generalized TSL, FLASSH 1.0 meets benchmark requirements, demonstrating noticeable agreement with experiment for both TSLs and the resulting integrated cross sections. Additional features including a graphical user interface (GUI), plotting diagnostics, and formatted output options including ACE files allow users to complete a TSL evaluation with minimal input and maximum flexibility. The user GUI creates input files for FLASSH, reducing user error and also providing built-in error checks. Autofill options and suggested input values help make TSL evaluation accessible to novice users. The FLASSH code is compiled to run on both Windows and Linux platforms with automatic parallelization. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

aphBO-2GP-3B: a budgeted asynchronous parallel multi-acquisition functions for constrained Bayesian optimization on high-performing computing architecture

High-fidelity complex engineering simulations are often predictive, but also computationally expensive and often require substantial computational efforts. The mitigation of computational burden is usually enabled through parallelism in high-performance cluster (HPC) architecture. Optimization problems associated with these applications is a challenging problem due to the high computational cost of the high-fidelity simulations. In this paper, an asynchronous parallel constrained Bayesian optimization method is proposed to efficiently solve the computationally expensive simulation-based optimization problems on the HPC platform, with a budgeted computational resource, where the maximum number of simulations is a constant. The advantage of this method are three-fold. Firstly, the efficiency of the Bayesian optimization is improved, where multiple input locations are evaluated parallel in an asynchronous manner to accelerate the optimization convergence with respect to physical runtime. This efficiency feature is further improved so that when each of the inputs is finished, another input is queried without waiting for the whole batch to complete. Second, the proposed method can handle both known and unknown constraints. Third, the proposed method samples several acquisition functions based on their rewards using a modified GP-Hedge scheme. The proposed framework is termed aphBO-2GP-3B, which means asynchronous parallel hedge Bayesian optimization with two Gaussian processes and three batches. The numerical performance of the proposed framework aphBO-2GP-3B is comprehensively benchmarked using 16 numerical examples, compared against other 6 parallel Bayesian optimization variants and 1 parallel Monte Carlo as a baseline, and demonstrated using two real-world high-fidelity expensive industrial applications. The first engineering application is based on finite element analysis (FEA) and the second one is based on computational fluid dynamics (CFD) simulations.

97 MATHEMATICS AND COMPUTING↗

A time-parallel multiple-shooting method for large-scale quantum optimal control

Quantum optimal control plays a crucial role in quantum computing by providing the interface between compiler and hardware. Solving the optimal control problem is particularly challenging for multi-qubit gates, due to the exponential growth in computational complexity with the system's dimensionality and the deterioration of optimization convergence. To ameliorate the computational complexity of time-integration, this paper introduces a multiple-shooting approach in which the time domain is divided into multiple windows and the intermediate states at window boundaries are treated as additional optimization variables. Further, this enables parallel computation of state evolution across time-windows, significantly accelerating objective function and gradient evaluations. Since the initial state matrix in each window is only guaranteed to be unitary upon convergence of the optimization algorithm, the conventional gate trace infidelity is replaced by a generalized infidelity that is convex for non-unitary state matrices. Continuity of the state across window boundaries is enforced by equality constraints. A quadratic penalty optimization method is used to solve the constrained optimal control problem, and an efficient adjoint technique is employed to calculate the gradients in each iteration. We demonstrate the effectiveness of the proposed method through numerical experiments on quantum Fourier transform gates in systems with 2, 3, and 4 qubits, noting a speedup of 80x for evaluating the gradient in the 4-qubit case, highlighting the method's potential for optimizing control pulses in multi-qubit quantum systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Decentralized Carrier Phase Shifting for Optimal Harmonic Minimization in Asymmetric Parallel-Connected Inverters

This paper presents a carrier phase shifting technique for minimizing the aggregate harmonics in networks of asymmetric parallel-connected inverters for distributed power generation system applications. The proposed technique is: 1) implemented in a decentralized manner, relying only on local voltage and current measurements, and 2) optimal in the sense that it minimizes a cost function representing the carrier-frequency current harmonics. The analysis indicates that the proposed optimal carrier phase shifting technique can enable order-of-magnitude reductions in harmonic power, and also universal improvements compared to symmetric carrier interleaving for asymmetric inverter networks. Moreover, compared to existing methods that require either centralized communication or information exchange between inverters to coordinate carriers, the proposed technique is completely decentralized, which provides important practical benefits for implementation, including improved robustness and reduced cost. The technique is experimentally validated on a network of three single-phase 2-kW inverters and demonstrates a 36.5% reduction in the weighted total harmonic distortion factor of the aggregate inverter current, and the ability to converge to the optimal carrier phase spacing dynamically in less than one line frequency cycle (16.7 ms) in steady state and transient operating conditions.

42 ENGINEERING↗

Parallel Solver Framework for Mixed-Integer PDE-Constrained Optimization

ROL-PEBBL is a C++, MPI-based parallel code for mixed-integer PDE-constrained optimization (MIPDECO). In these problems we wish to optimize (control, design, etc.) physical systems, which must obey the laws of physics, when some of the decision variables must take integer values. ROL-PEBBL combines a code to efficiently search over integer choices (PEBBL = Parallel Enumeration Branch-and-Bound Library) and a code for efficient nonlinear optimization, including PDE-constrained optimization (ROL = Rapid Optimization Library). In this report, we summarize the design of ROL-PEBBL and initial applications/results. For an artificial source-inversion problem, finding sources of pollution on a grid from sparse samples, ROL-PEBBLs solution for the nest grid gave the best optimization guarantee for any general solver that gives both a solution and a quality guarantee.

97 MATHEMATICS AND COMPUTING↗

Nonlinear multiobjective and dynamic real-time predictive optimization for optimal operation of baseload power plants under variable renewable energy

Considering the increase of disruptive variable renewable energy penetration into the power grid, this article focuses on the investigation of a multiobjective and dynamic real-time optimization framework to address the cycling of large-scale power plants under renewable penetration. In this framework, a parallelized particle swarm optimization step is first performed to generate feasible initial points. Then, a multiobjective and dynamic real-time optimization formulation generates optimal trajectories. Further, the benefit of predictive capability is investigated for the dynamic component, which introduces the novel nonlinear multiobjective and dynamic real-time predictive optimization approach. Two multiobjective formulations to obtain Pareto front optimal in real time are explored: the modified Tchebycheff-based weighted metric and ϵ-constraint methods. Economic and environmental objectives are considered in this study. A novel topical discussion on the intersection of dynamic real-time optimization with model predictive control is also presented. The developed framework is successfully applied to a baseload coal-fired power plant with postcombustion CO 2 capture. Results indicate that the approach can be deployed for a large-scale system if automatic differentiation, model reduction, and parallelization are adopted to improve computational tractability, with computational improvement up to 120-folds after performing these steps. Finally, market and carbon policies showed an impact on the optimal compromise between the objectives with an additional 63 ton of CO 2 captured under favorable market conditions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Scaling Automatic Vector Data Alignment to Satellite Imagery

Given the tremendous volume of accessible Earth Observation (EO) data, there is a need to develop scalable Geospatial Artificial Intelligence (GeoAI) solutions for time-sensitive applications. Scalability in this context refers to rapidly processing large-scale EO data using high performance computing resources. Accurate mapping of the built environment from remote sensing (RS) imagery has been one of the crucial components in GeoAI workflows for a wide spectrum of humanitarian applications. Derived vector data of built environment is often leveraged for disaster preparedness and response activities. However, factors such as differences in ortho-rectification, atmospheric conditions and human error, results in spatial misalignment between vector data and the timely available RS imagery. Model training for downstream tasks such as object detection, change analysis, etc., is negatively impacted due to such spatial misalignment. Although there has been progress towards automatic alignment of vector data, the lack of scalability remains an open research challenge. This paper proposes to leverage parallel computing to optimize an automatic vector data alignment workflow. It further employs CPU-level multi-core parallelism for improving the performance of the workflow for scalable built environment mapping. We report observations and discuss findings from the preliminary experiments performed on the Summit Supercomputer.

Potnis, Abhishek↗